Why LLMs Can't Hold a Patient Together

I once watched an AI model kill a patient.

The patient was simulated, a test case run through a multi-step scenario. You give a treatment, see how the patient responds, decide on the next treatment, and so on. It should have been a routine run. Three steps in, the model prescribed a drug the patient was allergic to.

What made it unsettling was that the model knew about the allergy. It had mentioned it itself, earlier in the same conversation. Somewhere between step one and step three, the fact just slipped away. And when the model made the fatal choice, there was no hesitation and no hedging. It was completely confident.

There was no one to blame. Just the design of the system doing what it was built to do.

If you've used ChatGPT or Claude, you may have seen a smaller version of this: you mention something early in a long conversation and it quietly gets lost later. In a chat about dinner recipes, that's annoying. In a model of a patient's body over a course of treatment, it's dangerous. A 2026 paper introducing EHRWorld, a model built specifically to track patients over time, describes the problem precisely: LLMs "struggle to maintain consistent patient states under sequential interventions, leading to error accumulation across multi-step interactions." The word I'd underline is accumulation. The mistakes compound. Each step makes the next one a bit worse.

People sometimes treat this as a prompting problem, as if the right instructions would make it go away. I don't think they will, and the reason is a little uncomfortable. For roughly a decade, a lot of AI research has worked on the assumption that intelligence is basically prediction. Large language models are the strongest evidence yet that it isn't. They're astonishing at predicting the next word. But knowing that a drug lowers blood pressure isn't a word-prediction problem. It's knowing how one thing causes another.

That's what researchers mean by a "world model": a representation of how things actually change. If I give this drug, what happens to the heart rate? If this lab value moves, what moves with it? Language models don't really have that. They have an extraordinary sense of what usually comes next in a sentence, and the next sentence is not the next state of a patient's body.

Self-driving cars taught the same lesson the expensive way. You can't drive a car with a model that describes roads beautifully. Medicine is that problem again.

So the answer isn't to wait for a smarter chatbot. It's to stop asking one tool to do two very different jobs. A 2025 paper in Health Policy breaks a digital twin into five parts: the patient, the data connection, the patient-in-silico (the computer version of them), the interface, and the synchronisation that keeps it all current. The authors argue that pairing AI with mechanistic modelling addresses "the limitations of either approach used independently." Mechanistic models are built from how the body actually works, with equations for how a drug moves through the blood or how the heart responds to stress.

In that setup, the language model does what it's good at: explaining, translating, talking with the clinician in plain language. The mechanistic model keeps track of the patient's actual state, and a simulation layer checks the work.

Now for the part where I'm guessing, and I want to be upfront that it's a guess. I think "patient state" is going to become infrastructure in its own right. Instead of being a pile of text in a chat window, it would be a permanent, versioned record that every action reads from and writes to, and that can't silently contradict itself. Banks have run like this for decades; your balance doesn't drift because a teller forgot a deposit. We haven't built it for medicine mostly because everyone has been staring at the models. If I had to bet, the most important company in this space ten years from now won't have the best AI. It will have built that state layer, the thing everyone else plugs into.

In that world, clinicians wouldn't really "prompt" anything. They'd say what they want in ordinary language, the system would turn it into a configured model, run it, and explain the result. The language model would be the translator, not the brain.

What actually worries me, more than any of the architecture, is who's building it. We're training AI people who've never studied physiology and clinicians who've never been taught how these models fail. Most of the real problems live in the gap between those two groups, and right now very few people are standing in it.

The patient's body doesn't care how elegant your system is. It does what it does. If your model can't keep track of that, it doesn't matter how good its next sentence sounds.

← Blog