Last week, I ended with a question: If we can change the temperature, what do we want the water to become?
I have been thinking about that question ever since.
In Part 1, I compared artificial intelligence to water, and our data, institutions, assumptions, and human systems create its container. AI learns within that container and, like water, it will always assume the shape of its container. And like water, which can freeze and become solidified in a container, when we automate what AI has learned and deploy it repeatedly across healthcare systems, we risk freezing some of those existing shapes into infrastructure.
But ice can melt.
That is where I want to begin today, because I am not interested in spending all of our time talking about what is wrong with artificial intelligence. I am much more interested in what becomes possible when we become intentional about what we teach it next.
If AI is learning from us, perhaps we should begin by asking who its teachers are. Not simply whether we need more engineers or larger datasets, but whether we have enough kinds of humans in the rooms where consequential decisions are made. We need people whose experiences cause them to ask different questions, people who recognize a blind spot because they have lived inside it, and those who notice who is missing because, at some point, they have been the person who was missing.
That is vital, particularly in healthcare.
Imagine an AI system designed to help clinicians make decisions about patients: Who decided what a “normal” patient looks like? Whose families were represented in the data? Were transgender patients included? Were there same-sex couples or people with disabilities? Were Black patients’ experiences of pain adequately represented? Were immigrants represented beyond assumptions about language, education, or socioeconomic status?
And when the system gets something wrong, who in the room recognizes it?
We already know these questions have consequences. In a landmark 2019 Science study, Ziad Obermeyer and colleagues examined an algorithm used to identify patients who might benefit from additional healthcare support. The algorithm used healthcare spending as a proxy for healthcare need. That choice sounds reasonable until we remember that what healthcare systems spend on a patient does not necessarily tell us how sick that patient is.
The researchers found that at the same algorithmic risk score, Black patients were considerably sicker than White patients. Because less money had historically been spent caring for Black patients with comparable needs, the algorithm interpreted lower spending as lower need. When the researchers corrected that disparity, the proportion of Black patients identified to receive additional help increased from 17.7% to 46.5%.
Think about what happened there. The algorithm did not need to be programmed to discriminate against Black patients. Race did not have to be in the instructions; in fact, race was specifically excluded from their instructions. But human beings chose a measurement, healthcare cost, from a system in which access and spending were already unequal. The algorithm assumed the shape of the container.
That finding also challenges one of the questions we commonly ask about artificial intelligence: Does it work? The algorithm could successfully predict healthcare costs. But healthcare cost was not the same thing as healthcare need. Perhaps, then, we need to ask something more fundamental: What exactly have we taught the AI to measure, and for whom does that measurement work?
A second study brings the concern even closer to clinical decision-making. In 2024, researchers writing in The Lancet Digital Health evaluated GPT-4, a generative artificial intelligence system built on what is called a large language model, a type of AI trained on enormous amounts of text to recognize patterns in language and generate responses. They tested it across medical education, diagnostic reasoning, clinical planning, and subjective patient assessment.
The researchers found that GPT-4 generated clinical cases that stereotyped demographic groups. When standardized patient scenarios were used, the differential diagnoses were more likely to include diagnoses associated with racial, ethnic, and gender stereotypes. Patient assessments and clinical plans also showed significant relationships between demographic characteristics and recommendations, including recommendations for more expensive procedures.
These are different studies involving different technologies, but together they point toward the same problem. Bias can enter through what we choose to measure, the information used to train a system, assumptions already embedded in language and institutions, and in how a tool interprets the human being standing in front of it.
The National Institute of Standards and Technology makes a similar distinction, identifying systemic, human, and computational or statistical sources of AI bias rather than treating bias as simply a problem of bad data.
So no, I don’t believe the solution is to fear artificial intelligence or somehow put it back into the box. I believe we have to become far more intentional about what happens before, during, and after we deploy it. We need to question what we call normal, examine the proxies we choose, test outcomes across different populations, listen when affected communities identify harms, and build accountability into the process rather than waiting for harm to become obvious.
Diversity at the end is an audit. Diversity at the beginning can change the design.
I have spent much of my career asking healthcare professionals to examine not only what they intend, but what patients actually experience. Intention → Impact. I believe we need to ask the same of artificial intelligence. That thinking has led me to a framework I call CAST™: Compassionate. Affirming. Safe. Trustworthy.
Is the AI Compassionate? I am not suggesting that a machine possesses compassion. I am asking whether the people designing and deploying it have considered the human being on the receiving end of its output. Does the system support care that preserves dignity and recognizes context, or does efficiency become more important than the person affected by the decision?
Is the AI Affirming? Does it make room for people as they actually are, rather than forcing them into assumptions about who they are supposed to be? If a system works beautifully for the presumed “default” patient but repeatedly misreads people whose race, gender, sexuality, disability, culture, or family structure falls outside that default, we cannot simply call that innovation and move on.
Is the AI Safe? Safety must mean more than whether the software functions as designed. We have to ask who could be harmed when it gets something wrong, whether those harms have been tested across populations, how quickly they will be recognized, and what happens when the same error is reproduced thousands or millions of times.
Is the AI Trustworthy? Can clinicians and patients understand enough about how consequential recommendations are being made to question them? Can the output be challenged? Is there meaningful human oversight? And when the system causes harm, is there accountability, or does responsibility disappear somewhere between the developer, institution, clinician, dataset, and machine?
Compassionate. Affirming. Safe. Trustworthy.
CAST™ is not intended to replace technical validation, regulation, clinical judgment, or rigorous bias testing. It asks us to add something that can disappear remarkably quickly when we become fascinated by what technology can do: the human consequences of what we are building.
And responsibility for those consequences cannot belong only to technologists. Physicians, patients, and ethicists all need to be in the room. People from historically marginalized communities need to be in the room, not simply after a product has been built so someone can ask whether it is “inclusive,” but early enough to influence what questions are asked, what gets measured, what gets tested, and what gets changed.
We cannot erase every human bias before building artificial intelligence. Human beings have been trying to understand our own biases for centuries. But we can become more intentional about what we automate. We can stop assuming that technologically sophisticated means equitable.
We can recognize that an algorithm can accurately answer the wrong question. And when someone says, “Wait. You forgot us,” we can make sure someone in the room has both the wisdom and the authority to stop, turn around, and listen.
Last week I wrote that AI didn’t invent our bias; it is learning it from us. This week, I want to add something else: bias isn’t the only thing it can learn from us. Compassion is human. Curiosity is human. Inclusion, courage, and accountability are human too. If our technology is learning from humanity, then humanity has a responsibility to become better teachers.
The water is already moving. We still have an opportunity to change the temperature, reshape the container, and decide more intentionally what becomes frozen into the healthcare systems of the future.
So perhaps the question isn’t simply, “What will artificial intelligence become?”
It is:
What are we C.A.S.T.ing into the future of healthcare?
With love and optimism,
Dr. Lulu®
P.S. Technology will keep learning. The question is whether we will.



