Musings · Essay

Reasoning from silence.

Why passing a medical exam is not the same as being fit to see patients.

Essay 12 min read August 2026

Large language models now pass medical licensing examinations. The best of them, tested on curated diagnostic cases, match or outperform attending physicians on structured reasoning tasks. This performance is real and remarkable, but it is being used to justify a step that the evidence does not support: deploying these systems to autonomously triage patients, with little or no clinician in the loop.

I have spent this year working with colleagues at Atman Labs Inc., Oxford and the University of Texas at Austin on a paper that sets out exactly why that step is premature. We published it on Arxiv this month. This is the version I would write if I were explaining it to someone who will never read the full preprint.

The exam is the wrong test

Every benchmark that shows an LLM performing at physician level shares a structural feature: the case is complete before the model sees it. A clinician assembled the history. The relevant details are present. The model is asked to reason over a finished document. In that setting, frontier models are genuinely impressive. OpenAI o1 recently outperformed attending physicians on structured emergency department cases at every diagnostic touchpoint measured. 94.9% of conditions were correctly identified. The headline writes itself.

Here is what the headline did not reveal: when real members of the public were given the same LLMs and asked to use them to navigate their own health problems, condition identification fell below 34.5%. The same model assessing the same disease. The same knowledge, applied to the same conditions, described by real people rather than assembled by clinicians. The performance collapsed by 60 percentage points because the input changed.

In the real world, the input is always a person who does not know which of the details matter.

What triage actually requires

Safe diagnostic triage requires sequential decision making under asymmetric cost. A missed subarachnoid haemorrhage and an unnecessary CT scan are not equivalent errors. They differ by orders of magnitude. The entire logic of clinical triage is built around that asymmetry: widen the differential when information is sparse, lower the threshold for escalation when a catastrophic diagnosis remains unexcluded, and keep asking until the dangerous possibility has been ruled out rather than simply unmentioned.

Consider the case we use in the paper. A 41-year-old woman opens a consumer triage tool and reports a headache that started the day before. She does not volunteer that the pain peaked within seconds, that it was the worst of her life, or that it began as she lifted a suitcase. The model does not ask. It produces a competent-looking answer: tension headache, migraine, and, prompted, subarachnoid haemorrhage among the possibilities. It advises that a lumbar puncture is not indicated. Nothing it was told was misinterpreted. The danger lay entirely in what it never asked.

Given a complete history, the same class of model identifies subarachnoid haemorrhage with 97.5% accuracy. The model knows about thunderclap headaches. It knows about lifting-triggered onset. It did not ask because asking, in the clinical sense of the word, is not what a model optimised to continue with the most probable text is built to do.

A model trained to extend on what is present has no native representation of what is absent. It conflates absence of evidence with absence of disease. A clinician trained to do the same thing could be dangerous. When LLMs behave this way in triage, they could be too.

Why the evidence has not exposed this

The failure is real. It is not evident in the published literature, the reason being structural rather than accidental.

The studies most often cited as proof of clinical readiness share a second feature alongside their curated inputs: they evaluate models on populations from which the difficult cases have already been removed. The most prominent conversational diagnostic AI, AMIE, demonstrated impressive performance in consultations with trained patient-actors over a text interface its own authors describe as unfamiliar in clinical practice, within an acknowledged limited scope of experimental simulated history-taking. A prospective feasibility study of the same system in a real clinic required patients to have already been determined by clinic triage staff not to need emergency care. The cases that define triage safety were removed before the system was engaged. A clinician supervised every consultation. The system never made the calls its results are taken to endorse.

This is not a criticism of individual studies. It is a description of how clinical AI evaluation currently works. Across 519 LLM evaluations published up to early 2024, only 5% used real patient care data. Across 4,609 studies to late 2025, fewer than 1 in 200 used real-world data in a prospective randomised design. The conditions that constitute triage risk, the rare, the atypical, the undifferentiated, the incompletely disclosed, are systematically removed before any model is scored. An evaluation that removes the catastrophic tail cannot detect a model that mishandles it.

And silence is being interpreted as reassurance. The field has measured what these systems do with the patients it selected for them. It has not measured what they do with the patient who walks in frightened, half-able to say what is wrong, and not knowing which detail changes everything.

The behaviours that compound the problem

The core failure, what we call reasoning from silence, does not act alone. It is compounded by a cluster of properties that make a good assistant and a dangerous triage system.

The first is credulity. When fabricated clinical details are embedded in a prompt, models often fail to flag the inconsistency and instead elaborate clinical reasoning around the false premise, in 50% to 82% of cases across models and mitigation strategies. On a benchmark of misleading medical context, accuracy fell from 71% on clean questions to 38% once a plausible falsehood was introduced. A safe history-taker probes an implausible claim. A fluent assistant accepts it as material to build on.

The second is agreeableness. Models tuned to be warmer and more empathetic show higher rates of affirming a user's incorrect belief, particularly when the user expresses distress. Triage often begins with minimisation: it is probably just stress, I do not want to waste anyone's time. Safe triage requires the system to resist reassurance when a high-harm diagnosis remains unexcluded. A model optimised for agreeable interaction may align with the patient's minimisation instead.

The third is miscalibration. In direct tests of judgement revision as new information arrives, leading models were systematically overconfident and rarely selected a neutral response. A system whose expressed certainty is not calibrated to what is missing cannot safely use confidence as a disposition signal.

None of these properties require a model to hallucinate to cause harm. In a benchmark of 4,249 real management decisions, 76.6% of errors were omissions: failures to recommend an action that was needed. The system need not invent a dangerous treatment. It can fail simply by accepting the incomplete history, confirming the minimisation, and closing the case before the dangerous possibility was excluded.

What would demonstrate readiness

If the problem were knowledge, the remedy would be a better model. The deficit is at two levels: the objective an LLM optimises, and the evaluation by which we decide when it is ready to be used clinically. Both must be addressed. The second is actionable now.

An adequate evaluation of autonomous triage must do 3 things that current practice does not. First, withhold information by design: generate patients whose histories are deliberately incomplete, with the ground-truth diagnosis known to the evaluator but not the model, because only by controlling what the patient does not say can an evaluation measure whether a system reasons from silence or merely completes the probable. Second, score the right objective: top-k diagnostic accuracy is silent on safety; a fit-for-purpose evaluation must reward widening under uncertainty, lowering the threshold for escalation, deferring when information is insufficient and resisting minimised histories, with the asymmetric cost of error written into the metric itself. Third, validate any simulator before trusting it: simulator realism must be demonstrated across face validity, distributional validity and predictive validity, evidence that performance on the simulator actually forecasts performance on real patients with real outcomes. Until that holds, strong simulator performance is a screening hurdle, not a warrant for deployment.

The claim requires the evidence

The case for autonomous LLM triage currently rests on evidence that does not measure the task most relevant to safety. Examination performance demonstrates medical knowledge. It does not show whether a system can manage danger under incomplete and ambiguous information, where the decisive fact may be one the patient has not volunteered.

The claim that an LLM can be trusted to triage the undifferentiated patient without a clinician in the loop is among the strongest claims in medicine. The evidence currently marshalled in its favour is almost entirely from settings designed to exclude the failure it does not see.

Closing that gap is not optional. For autonomous triage, demonstrated safety under realistic incomplete-information conditions is the precondition for everything that follows. Benchmark accuracy and exam scores are not enough. Safety should be demonstrated when the patient has not yet said the thing that matters.

← All musings The Evidence Constraint →