Can AI evaluate AI researchers? New study puts autonomous frameworks to the test
A new study evaluates autonomous AI researchers using state-of-the-art models such as GPT-5.4. Results show that reference papers from FARS outperform leading AI frameworks in scientific quality.

What happened?
A new research study presents a standardized benchmark for evaluating AI systems that autonomously conduct research and author scientific papers. The researchers compared four leading frameworks for autonomous AI researchers—Sakana AI (versions 1 and 2), CycleResearcher, and Data-to-Paper—against reference papers from the company FARS. The evaluation was conducted via an automated peer-review process powered by state-of-the-art models such as GPT-5.4, Gemini, and Claude across four dimensions: originality, scientific rigour, clarity, and significance.
Key facts
Why it matters
The results show that FARS's own reference papers significantly outperformed all four of the autonomous AI frameworks evaluated. This highlights a significant gap between what current autonomous research agents can produce and the quality required for high-level scientific review. The study establishes a necessary methodology for measurably tracking progress in automatically generated research.
Who is affected?
The study is of particular interest to AI researchers, academic journals, and technical organisations exploring the automation of scientific discovery. It also affects developers of large language models who seek to understand the limitations of current AI agents regarding complex logic and research methodology.
Impact on the EU
The research highlights challenges directly linked to EU standards for transparency and scientific quality in AI. As AI-generated research becomes more common, regulations such as the EU AI Act may impose higher requirements for the traceability and evaluation of automatically generated conclusions.
What else you should know
The study was conducted using a strict evaluation model based on three independent state-of-the-art models: GPT-5.4, Gemini, and Claude. All four frameworks were run against the same set of 15 research proposals from FARS, resulting in 60 generated papers that were directly compared with FARS's own 15 reference papers.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka AI-modeller användes för att granska artiklarna?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Can AI evaluate AI researchers? New study puts autonomous fr"