New research reveals flaws in evaluation of visual AI models
A new study identifies several methodological issues distorting how the field measures visual-linguistic reasoning capabilities in AI models, specifically within physics.

What happened?
Researchers have conducted an audit of the current evaluation chain for multimodal physics models. The study, titled "Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning" and published on 26 May 2026, points to three main problems: contamination between training and evaluation data, drift in translations, and an over-representation of multiple-choice questions in evaluation datasets. These factors lead to an inaccurate representation of the AI models' true capabilities.
Key facts
| Publikationsdatum | 26 maj 2026 |
|---|---|
| Antal brister identifierade | 3 |
| Kontamination (SciInstruct) | 134 nära-dubletter, 4 846 parafraser |
| Sonnet 4.5 prestandaskillnad | 30.5% vs 13.6% |
”We audit the multimodal-physics evaluation pipeline end-to-end and document three undetected construction practices that distort how the field measures vision-language reasoning: train-eval contamination, translation drift, and MCQ saturation.”
”A three-stage audit [...] surfaces 134 near-duplicates and 4,846 paraphrase candidates in SciInstruct alone.”
”A 46-pp format-and-novelty gradient on identical Sonnet weights between MCQ (79.7% on PhyX) and open-ended olympiad evaluation (33.4% on PhysOlym-A).”
Why it matters
These identified flaws mean current evaluation methods may exaggerate AI models' ability to perform visual-linguistic reasoning. Data contamination suggests models are inadvertently trained on parts of test sets, while translation drift and the focus on multiple-choice questions fail to reflect the complexity of real-world problem-solving. This hinders the accurate assessment of progress and the identification of truly intelligent systems.
Who is affected?
AI researchers and developers, particularly those working with multimodal models and visual-linguistic tasks, are directly affected. Organisations evaluating AI systems for complex applications, such as in education or science, must consider these findings to ensure the validity of their assessments. Users of AI models should also be aware of the methodological challenges in the evaluation process.
What else you should know
As part of the study, four new artefacts, PhysCorp-A, are being released to address the identified flaws and provide improved tools for evaluation.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka bolag berörs?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "New research reveals flaws in evaluation of visual AI models"