Nuanced clinical reasoning remains beyond the ken of frontier LLMs
When used straight off the shelf, even the most advanced large language models are unfit for use in advanced clinical decision-making.
That’s not to say these frontier models are of no use to clinicians. In fact, they are impressively accurate when making final diagnoses.
The rub is that they’re lousy at differentiating between possible diagnoses when several are in the running.
The findings are from a study conducted at Harvard and published online April 13 in JAMA Network Open.
Senior study author Marc Succi, MD, and colleagues designed the research to cut through the din of LLM vendors’ claims of clinical fitness.
Throughout 2025, the team put 21 off-the-shelf LLMs through their paces on the tracks of 29 clinical vignettes.
The tested models were considered “state of the art” during the study period. They included the latest iterations of ChatGPT-5 (from OpenAI), Claude 4.5 Opus (Anthropic), Gemini 3.0 Flash and Pro (Google), and Grok 4 (xAI).
One of the measures Succi and co-researchers used was a scoring system they developed.
Called the Proportional Index of Medical Evaluation for LLMs (PrIME-LLM), the novel system trialed the models across five domains of clinical reasoning. These were differential diagnosis, diagnostic testing, final diagnosis, management and miscellaneous clinical reasoning questions.
The team also pitted models against one another, comparing performances by such metrics as overall showing and demographic breakouts.
Imaging inputs improve LLM accuracy
The researchers’ key findings included:
- PrIME-LLM scores ranged from 0.64 (Gemini 1.5 Flash) to 0.78 (Grok 4), with reasoning-optimized models outperforming nonreasoning models and GPT models scoring highest overall.
- Differential diagnosis was less accurate than diagnostic testing, while final diagnosis, management and miscellaneous reasoning were more accurate.
- Failure rates exceeded 0.80 for differential diagnosis in all models but were less than 0.40 for final diagnosis.
- Multimodal performance was robust; most LLM models showed improved accuracy with image inputs.
In their discussion section, Succi and co-authors comment that the promise of LLMs in clinical medicine lies in their potential to augment—not replace—physician reasoning.
“This study establishes the first benchmark for longitudinal clinical reasoning (to our knowledge) and introduces the PrIME-LLM framework, a multidimensional metric designed to capture performance across the full arc of diagnostic and management tasks,” they write.
Dr. AI deserves its reputation for uneven expertise
The risk that the study clarifies, the authors add, “is not just that LLMs are sometimes wrong but that their reasoning is brittle precisely where uncertainty and nuance matter most.”
“Benchmarks that reward only correct final answers risk reinforcing [diagnostic] shortcutting, widening the gap between marketing claims and the skills actually required at the bedside,” they underscore.
More:
As commercial systems increasingly market reasoning capabilities and move toward clinical deployment, PrIME-LLM scores provide an independent, reproducible and extensible benchmark to track progress, expose persistent limitations and guide safe integration into healthcare practice.
The findings are available in full for free.
