Magis-Bench: New benchmark tests AI in judicial tasks
A new benchmark, Magis-Bench, evaluates the ability of large language models (LLMs) to perform judge-related tasks. It focuses on assessing legal arguments, applying doctrine to facts, and making reasoned decisions.

What happened?
Magis-Bench consists of 74 questions from eight Brazilian judicial exams conducted between 2023 and 2025. These include tasks requiring legal analysis as well as practical exercises such as formulating complete civil and criminal judgements. A total of 23 state-of-the-art LLMs were evaluated using an "LLM-as-a-judge" methodology, where four independent models acted as assessors.
Key facts
| Antal frågor i Magis-Bench | 74 |
|---|---|
| Tidsperiod för examensprov | 2023-2025 |
| Antal utvärderade LLM:er | 23 |
| Antal bedömande modeller | 4 |
”Existing benchmarks for legal AI focus primarily on tasks where LLMs must produce legal arguments or documents, yet the capacity to judge such arguments—weighing competing claims, applying doctrine to facts, and rendering reasoned decisions—is arguably as fundamental to a well-fu”
Why it matters
While existing AI benchmarks in the legal field primarily focus on generating legal arguments or documents, the ability to assess such arguments is central. Magis-Bench fills a gap by measuring the capacity of LLMs to weigh conflicting claims and apply legal doctrine. This provides insight into how well AI can handle complex decision-making processes within the justice system.
Who is affected?
Developers of large language models are affected as Magis-Bench offers a new method to evaluate AI's legal competence. Legal professionals and justice systems globally can potentially see how AI might complement or enhance judicial review. Users of legal AI tools can expect future refinements in systems that better understand and make legal decisions.
What else you should know
The results of the evaluation show a high level of agreement between the assessing models (Kendall's tau).
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka typer av uppgifter ingår i Magis-Bench?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Magis-Bench: New benchmark tests AI in judicial tasks"