Doctoral researcher Marta Arbizu Gómez analyzes the current scientific evidence on the limitations of language models in clinical reasoning and their impact on everyday clinical practice.
A study published in JAMA Network Open reveals that large language models (LLMs) have an error rate above 80% in differential diagnosis, despite their final-answer accuracy. In neuropsychological assessment and neurorehabilitation, these data show that AI does not replace human clinical reasoning and should be used only as a complementary tool under professional supervision.
Why is it important to assess the clinical reasoning of artificial intelligence?
Large language models, known as LLMs, are playing an increasingly prominent role in health care. They are used or being explored to summarize medical records, support decision-making, answer medical questions, generate reports, and assist with clinical documentation.
However, answering an isolated medical question correctly is very different from reproducing the complete clinical reasoning process carried out by a healthcare professional when evaluating a real patient. In clinical practice, a diagnosis does not usually arise from a single answer, but from a gradual process: information is gathered, hypotheses are formulated, tests are ordered, possibilities are ruled out, and a management plan is decided.
Until now, many evaluations of artificial intelligence in medicine have been based on multiple-choice examinations, such as medical licensing tests or multiple-choice questions. Although these formats are useful for measuring knowledge, they do not fully capture the complexity of clinical reasoning, especially in situations of uncertainty.
The study published in JAMA Network Open addresses precisely this question: can current LLMs reason reliably throughout the entire clinical workflow?

Subscribe
to our
Newsletter
How was the research on LLMs in medicine conducted?
The authors evaluated the performance of 21 state-of-the-art language models using 29 standardized clinical vignettes from the MSD Manual. These vignettes simulate clinical cases step by step, from the patient’s initial presentation through diagnosis and management.
The models evaluated included systems from different families, such as GPT, Claude, Gemini, DeepSeek, and Grok. They included recent models optimized for reasoning, such as GPT-5, Claude 4.5 Opus, Gemini 3.0 Pro, and Grok 4.
Each model had to answer questions organized into five domains of clinical reasoning:
- Differential diagnosis;
- diagnostic testing;
- final diagnosis;
- management or treatment;
- and additional clinical questions.
To evaluate performance, the researchers did not limit themselves to calculating overall accuracy. They introduced a new metric called PrIME-LLM, which summarizes the model’s balanced performance across five areas of clinical reasoning. This metric is represented using radar charts: the broader and more balanced the chart area, the better the model performs throughout the clinical process.

What do the key results reveal about LLMs in diagnosis?
The results show a very clear pattern: current LLMs can achieve good results when asked to identify the final diagnosis, but they have far greater difficulty when asked to generate a differential diagnosis.
This is especially relevant because differential diagnosis is a central part of medical reasoning. It involves considering different diagnostic possibilities, maintaining uncertainty, and avoiding premature closure. According to the authors, models tend to converge prematurely on a single answer instead of reasoning more flexibly, as a clinician would.
Models optimized for reasoning achieved better results than non-optimized models, but this improvement did not eliminate the main limitations. Even the most advanced models continued to show important failures in the early stages of clinical reasoning.
Overall, the best PrIME-LLM results were achieved by models such as Grok 4, GPT-5, GPT-4.5, Claude 4.5 Opus, Gemini 3.0 Flash, and Gemini 3.0 Pro. However, the study warns that high overall accuracy may conceal important weaknesses if a model does not perform in a balanced manner across all tasks.
| Metric or result | Key finding |
|---|---|
| Models evaluated | 21 LLMs |
| Clinical cases used | 29 clinical vignettes |
| Total responses analyzed | 16,254 |
| Best overall performance | Recent models optimized for reasoning |
| Strongest area | Final diagnosis |
| Weakest area | Differential diagnosis |
| Failure rate in differential diagnosis | Above 0.80 across all models |
| Usefulness of PrIME-LLM | Detects imbalances that overall accuracy may conceal |
| Clinical conclusion | LLMs are not safe for autonomous use without supervision |
What are the implications for clinical practice?
This study raises an important warning: it is not enough for a model to get final diagnoses right to consider it safe for use in medicine.
In clinical practice, one of the greatest risks is that a tool may generate an apparently correct answer without adequately considering other possibilities. This can create a false sense of security, especially in highly uncertain contexts.
The results suggest that LLMs may be useful for specific clinical tasks, especially when those tasks are clearly defined and supervised by professionals. For example, they could help summarize information, structure data, support documentation review, or serve as a complementary tool in low-uncertainty settings.
However, the authors are clear: current models should not be used as autonomous diagnostic agents or in patient-facing contexts without clinical supervision. Their performance still does not reproduce the complexity of human medical reasoning, especially during the early stages of diagnosis.
How does this advance relate to NeuronUP?
At NeuronUP we work with evidence-based digital tools for cognitive rehabilitation and cognitive stimulation. In this context, studies like this are especially relevant because they help define the realistic role of artificial intelligence in health care.
AI can add value when integrated responsibly, transparently, and under supervision. For example, it can help:
- Support the organization of clinical and functional information.
- Facilitate structured follow-up of patients.
- Help personalize cognitive interventions based on objective data.
- Complement, but not replace, professional judgment.
This article reinforces a key idea: health technology must be evaluated not only by its ability to provide correct answers, but also by its safety, reliability, and usefulness within the complete clinical process.
In neurorehabilitation, this means moving toward models in which artificial intelligence can support decision-making, but always as part of professional teams and guided by clear clinical criteria.
Conclusion
The study shows that language models have made remarkable progress and can achieve good results on specific medical tasks. However, they continue to have important limitations in longitudinal clinical reasoning, especially when generating differential diagnoses and managing uncertainty.
The new PrIME-LLM metric makes it possible to evaluate these models more comprehensively than overall accuracy, as it identifies imbalances across different stages of the clinical process.
Ultimately, LLMs may be promising tools for supporting certain health care processes, but they are not yet ready to replace human clinical reasoning or make diagnostic decisions autonomously.
References
- Rao AS, Esmail KP, Lee RS, Jiang S, Arraiza Carlo B, Gill J, Khanna P, Kalmowitz E, Montagnese B, Heydari K, Jiao Q, Bott E, Nguyen D, Wang G, Hood M, Landman AB, Succi MD. Large Language Model Performance and Clinical Reasoning Tasks. JAMA Network Open. 2026;9(4):e264003. doi:10.1001/jamanetworkopen.2026.4003.
Frequently asked questions about LLMs in medicine
1. Can language models (LLMs) perform clinical diagnoses autonomously?
No. The models evaluated had a failure rate above 80% when generating differential diagnoses and showed a tendency to converge prematurely on a single option. Therefore, they are not safe to operate autonomously without direct supervision from a healthcare professional.
2. What is the PrIME-LLM metric, and why is it important?
It is a metric designed to assess the models’ balanced performance across five domains of the clinical process: differential diagnosis, testing, final diagnosis, treatment, and additional questions. Its importance lies in detecting weaknesses and imbalances in reasoning that traditional overall-accuracy metrics often conceal.
3. What are the main limitations of artificial intelligence in the clinical workflow?
Although advanced model families such as GPT, Claude, Gemini, Grok, and DeepSeek excel at identifying the final diagnosis, they experience a severe decline in performance during the initial stages. Their greatest limitation is the inability to manage diagnostic uncertainty and weigh different hypotheses before closing a case.
4. What implications do these findings have for cognitive assessment and rehabilitation?
The study demonstrates that AI should not replace professional judgment in neuropsychological assessment. However, technology remains an excellent complementary resource for organizing clinical information, personalizing cognitive interventions, and optimizing patient documentation under specialist guidance.







A Guide for Professionals: Addressing Dyslexia in Children with ADHD Through Evidence-Based Neuropsychology
Leave a Reply