Specialised small language models outperform giants in law
A study shows that specialised, domain-trained Small Language Models (SLMs) outperform Large Language Models (LLMs) in legal contract data extraction, with significantly lower costs and higher precision.

What happened?
A new study published on arXiv compared the performance of Olava Extract, a self-hosted SLM, with five advanced LLMs for structured data extraction from legal agreements. The results show that Olava Extract achieved the best overall performance with a macro F1 score of 0.812 and a micro F1 score of 0.842. Furthermore, Olava Extract reduced operating costs by 78% to 97% compared to the LLMs tested.
Key facts
| Publikationsdatum | 7 maj 2024 |
|---|---|
| Bästa Macro F1-värde | 0.812 (Olava Extract) |
| Bästa Micro F1-värde | 0.842 (Olava Extract) |
| Kostnadsreduktion | 78% till 97% |
”Olava Extract achieved the strongest aggregate performance in the study, with a macro F1 of 0.812 and a micro F1 of 0.842, while reducing inference cost by 78% to 97% compared with the frontier models tested.”
”The findings shows that high performing, human comparable legal AI no longer requires the largest externally hosted models.”
Why it matters
This development is significant as it challenges the perception that commercially valuable enterprise AI requires the largest and most expensive models. The study indicates that smaller, domain-specific models can deliver superior performance in niche applications, particularly where precision and the minimisation of "hallucinations" are critical, such as in the legal field.
Who is affected?
The results primarily affect developers and firms within law and finance that use or are considering implementing AI for contract review. They demonstrate that investing in specialised solutions can be more effective than relying on generic, large models. AI researchers can also benefit from these insights into the potential of SLMs.
What else you should know
The study underscores the importance of precision in legal AI applications, where incorrect extractions can lead to operational risks and increased manual review.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka bolag berörs?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Specialised small language models outperform giants in law"