Now is the Time to Build Your AI Governance Program
Healthcare Tech Outlook

A featured contribution from Leadership Perspectives: a curated forum reserved for leaders nominated by our subscribers and vetted by our Healthcare Tech Outlook Advisory Board.

Marshfield Clinic Health System

Now is the Time to Build Your AI Governance Program

Mitch Kwiatkowski

Look at any trade journal, social media feed, or news outlet, and you will read that generative AI will become a pervasive force in healthcare. Large Language Models (LLM) promise to create new insights that healthcare organization can leverage in countless ways, from clinical decision support to diagnostic discovery to operational optimization. This is great news for patients, but it does not come without risks.

Artificial Intelligence is not inherently good or bad. Someone creates it and determines what it will do. A person chooses to invoke agency with the help of AI, or they relinquish choice in deference to AI’s autonomous decisions. In any stage of development and use, we have a responsibility to understand an AI solution, observe how it is used, assess the impact it has on people, and respond when it causes unintended consequences. Unfortunately, many AI solutions that are created or purchased are ‘black boxes.’ We do not know an algorithm's inner workings to understand how it makes recommendations and decisions. In cases where documentation is provided, it may not be presented in a way that the average person can understand. These can hide risks of potential harm.

Examples of AI Risks

There are plenty of recent examples demonstrating the risks of AI. In 2019, Science published an article that concluded a commercial algorithm used by many hospital systems across the U.S. was tainted by racial bias because it used cost as a predictor. The model incorrectly concluded that because less money was spent, black patients were healthier than white patients in that same category.

A NIHCM study in 2021 described several more examples of racial bias in healthcare. Algorithms contributed to 4.7 times more racial disparities when measuring pain compared to what a radiologist would score manually. An algorithm developed by the American Heart Association assigned more points to non-black races which labeled black patients as lower risk of death. A common calculation used to score kidney function would often result in higher results for black patients, leading to a false conclusion of better kidney health and a delay in care.

Bias is not limited to race or ethnicity. A 2016 study by the NHS in England found that 50 percent of women were misdiagnosed after a heart attack. Symptoms are typically different from men, and algorithms used to support physicians at the time did not take those differences into account. Other studies have found that the treatment of obese patients could be affected by bias. Obesity is often linked with complex chronic conditions, and the failure of a model to account for those might oversimplify (or overlook) recommendations or predictions of compliance to prescribed regimens. A model might also incorrectly link the cause of a condition to obesity rather than another (potentially more serious) cause.

An emerging trend in analytics is the use of low-code/no-code tools to create AI algorithms. The age of the ‘citizen data scientist’ sounds like a panacea for most industries, where the average person can create predictive models without knowledge of statistics or programming. However, access to self-service does not mean the quality of what gets produced is on par with a professional data scientist.

In the world of healthcare, our decisions can have life-or-death consequences. The ability to create models quickly by the average person does not mean that is the right thing to do for patients. 

A Practical Example

To best demonstrate what an organization could face, consider the following example from a paper I wrote last year. In that project, I used an off-the-shelf low-code/no-code tool and publicly available data to predict the likelihood that a diabetic patient would be readmitted to the hospital within 30 days of discharge. The tool processed thousands of rows of demographic, clinical, and event data and produced a list of the best predictive models in under an hour.  

On the surface, the power of the best model was compelling. Performance was optimal, and it predicted true readmissions 81.6 percent of the time. Results were further validated with two other separate sets of de-identified data that contained the same attributes as the training file. For many organizations, implementation of the model might be a no-brainer. Our ability to predict 30-day readmissions could lead to improved clinical outcomes, lower costs, and increased patient satisfaction.  

But the purpose of my project was not to find the best model. It was to see what underlying elements of bias or fairness might exist. On deeper analysis, a glaring issue of bias existed in the model that resulted in unfair results. Despite its relatively high accuracy, black patients were 21.8 percent more likely to be incorrectly labeled as not at risk for readmission than caucasians. An organization using this model as built would overlook opportunities to help black patients who are truly in need of intervention to prevent readmission to the hospital. We could also infer that caucasian patients who are false positives – predicted to be readmitted – would receive attention that they did not truly need.

Tweaking the predictive features and the parameters of the model maintained the same level of accuracy but only reduced the biased impact of these results by 2.1 percent.

"Understanding AI will ensure we maximize the value it provides to patients and reduces the chance that we inflict harm or create unintended consequences"

Of course, the model could be implemented as it is. Given its high level of accuracy, an argument could be made that predicting some readmissions is better than none. But is that the right thing to do? Are we treating patients fairly by implementing a model that we now know is biased and unfair for some of the population?   

These are the types of challenges we have to consider when using AI technologies. The impact of our decisions will range from small to life-threatening, but it is our duty as healthcare professionals to identify and mitigate the risks at any level.

A Practical Approach to AI Governance

Even with examples of misuse, bias, and unfair AI practices in the real world, leaders may not know where to begin. Using resources already available, it is possible to establish the foundation of an AI governance framework that can be scaled and expanded across the enterprise. This does not require a financial investment to start, but it requires focus, time, and commitment. 

Leadership

The core of a robust governance framework starts with a strong sponsor. This can come from any leader in the organization like the head of data and analytics or someone in the risk or compliance department. The sponsor will help define ‘artificial intelligence’ for the organization and leverage relationships with other leaders to communicate the importance of the program and the value it brings to patients. Like many risk and compliance topics, the value of AI governance may be in preventing what might happen versus avoiding what will happen.

Organization

Once organizational leadership supports the initiative, the next step is to establish a governance committee or council comprised of diverse perspectives from the enterprise including compliance, clinical, technical, and operational roles. For each AI solution, the committee will have to review and understand the business case, how a model works, how a model’s results will be used, and who will use them. In large or complex organizations where a single committee is untenable, subcommittees and workgroups can isolate focus on specific business needs.

Process

Each request to acquire or develop AI technology should follow a robust, agile process that includes intake, deliberation, monitoring, and mitigation. Documentation should include the intent of the solution, who will use it, who or what it will apply to, and what steps have been taken to mitigate any potential harms such as bias, fairness, or unintended consequences. Institutional Review Board (IRB) processes in research institutes are often collecting this type of information, so they may be good places to start for AI-specific documentation. Once a solution is deployed to production, it should be audited regularly. If the model drifts or unintended consequences are discovered, action should be taken to quickly remediate.

Commitment

Some conversations during the AI governance process will be difficult, and people might get frustrated with more red tape. Rather than discussing issues that have simple yes-or-no answers, organizations may have to talk through ethical dilemmas that will test their ability to adhere to their mission, vision, and principles. Choosing a path for the sake of competitive advantage (i.e., ‘our competitors will do it, so we have to.’) or innovation may seem like an easy answer, but easy is not always right. Tough questions and courageous conversations are necessary to highlight the potential risks and unintended consequences of AI. The goal of AI governance is not to say no; its purpose is to ensure AI solutions are understood and any actions taken with or by them are defensible. As everyone gets comfortable with the process, it can be optimized using precedents and expedited review.

Conclusion

Artificial intelligence regulations are taking shape and some like the EU AI Act could be right around the corner. These are generally designed to provide rules at a high level for AI vendors and technology companies, so they may not be enough to manage practical, day-to-day AI needs within a healthcare organization. However, these can be good guides for a program.  For example, using the EU AI Act’s definition of artificial intelligence or its risk profile approach can set the groundwork for what is in scope of a governance program.

Responsible AI requires attention, discipline, and effort. The time to establish a governance program is now. AI solutions that are transparent and explainable will build trust and confidence, and a strong enterprise program will guide the organization on a more equitable, inclusive, and ethical path. Understanding AI will ensure we maximize the value it provides to patients and reduces the chance that we inflict harm or create unintended consequences.

The articles from these contributors are based on their personal expertise and viewpoints, and do not necessarily reflect the opinions of their employers or affiliated organizations.

Weekly Brief