Study reveals failing AI judges in SQL systems – Qwen saves millions
A new study reveals that AI models used as judges in production systems can have unexpectedly low accuracy. By switching models, researchers both increased precision significantly and drastically cut costs.

What happened?
Researchers have in a new study on arXiv (2609.30290) examined the performance of the evaluation model gpt-4o-mini in a live deployed text-to-SQL pipeline. The results show that gpt-4o-mini performed extremely poorly against human annotators, with a Cohen's kappa value of a mere 0.04 in a targeted sample and 0.42 in a random sample. The model incorrectly flagged 77.1 percent of correct, human-verified answers due to so-called grade-hallucination.
Key facts
| gpt-4o-mini Cohen's kappa (berikat) | 0,04 |
|---|---|
| gpt-4o-mini felaktigt flaggade fall | 77,1 % |
| Qwen3.6-27B Cohen's kappa | 0,72 |
| Claude Opus 4.7 Cohen's kappa | 0,71 |
| Kostnadsminskning med Qwen3.6-27B | Ca 1/300 av kostnaden |
Why it matters
The study shows that using cheaper LLM models as automated judges can lead to systematic errors in production environments. By replacing gpt-4o-mini with the self-hosted model Qwen3.6-27B, the agreement (kappa) increased to 0.72, which is on par with Claude Opus 4.7 (0.71). At the same time, the operating cost per call was reduced to approximately one-three-hundredth of the previous cost.
Who is affected?
The study primarily concerns AI developers, data engineers, and companies using LLM-as-a-judge for automated quality assurance in production environments. The results demonstrate that smaller or cheaper models from leading providers are not always sufficient for critical evaluation without rigorous auditing.
What else you should know
The researchers noted that the error was exacerbated when weak evaluators were combined with stronger ones in an ensemble, which actually lowered overall accuracy. It was only when three strong models were used with a requirement for unanimity that Cohen's kappa rose to 0.79 with nearly 90 percent automatic coverage.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka berörs av detta?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Study reveals failing AI judges in SQL systems – Qwen saves "