Study: Language affects medical diagnostics in large language models
A new study shows that the diagnostic capabilities of large language models in medicine vary depending on whether they are prompted in English or French, with generally lower performance for French.

What happened?
Researchers investigated how five large language models (LLMs) perform in medical diagnostics when prompted in English versus French. The study was based on 180 clinical vignettes across 16 medical specialities. Two physicians assessed the models' diagnostic accuracy and reasoning quality using an 18-point scale.
Key facts
| Publikationsdatum | 19 maj 2026 |
|---|---|
| Antal modeller testade | 5 |
| Antal kliniska vinjetter | 180 |
| Antal medicinska specialiteter | 16 |
| Modeller presterade bättre på engelska | 4 av 5 |
| Medelskillnad i prestanda | 0.37-0.91 |
”Prompting language influences diagnostic reasoning and accuracy of large language models.”
”Four of the five models performed better in English (mean difference 0.37-0.91, adjusted p < 0.05), with the gap spanning multiple aspects of reasoning, including differential diagnosis, logical structure, and internal validity.”
”o3 was the only model showing no overall language effect.”
Why it matters
The results indicate that the choice of language in prompting can have a significant impact on the ability of LLMs to perform complex medical tasks. Dissimilarities were observed across several aspects of medical reasoning, including differential diagnoses, logical structure, and internal validity. This highlights the importance of multilingual evaluation for applications within clinical decision support.
Who is affected?
The study is relevant for developers of large language models, particularly those intending to implement them in global healthcare systems. Healthcare professionals and policymakers considering the use of AI for clinical decision support are affected, as language-specific performance differences can impact patient safety. Future users of AI-based diagnostic tools will also be affected.
Impact on the EU
The study is relevant for the EU, where many different languages are used within healthcare. Performance differences between languages could affect the implementation of AI tools for clinical decision support in member states. The availability of safe and effective AI solutions across language borders is a key issue for the EU's digital agenda.
What else you should know
One of the five models, o3, showed no overall language impact in the study. The research underscores the need for comprehensive testing across multiple languages before medical LLMs are widely implemented.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka modeller berördes?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Study: Language affects medical diagnostics in large languag"