Study compares performance of LLMs and fine-tuned models in NVDRS
A new study published on arXiv examines how Large Language Models (LLMs) and fine-tuned models perform in extracting circumstances from death investigations within the framework of the National Violent Death Reporting System (NVDRS) in the United States.

What happened?
Researchers have developed an algorithm, called 'Complexity Score', to predict when detailed prompts with full coding guidelines yield better results than simpler name-based prompts. A hybrid method has been constructed that selects the prompting strategy per circumstance. The study evaluated the performance of LLMs like GPT-5.2, Gemini 2.5 Pro, and Llama-3 70B against fine-tuned RoBERTa models.
Key facts
”We found that LLMs substantially outperform on low-prevalence circumstances where training data is insufficient. We further demonstrate that our framework generalizes across frontier LLMs, with GPT-5.2, Gemini 2.5 Pro and Llama-3 70B showing consistent performance benefits.”
Why it matters
The study highlights the challenge of extracting structured information from narrative texts, especially when semantic inference is required beyond simple keyword matching. The results show that LLMs perform significantly better in low-prevalence circumstances where training data is insufficient for fine-tuning. This is important for public health work regarding suicide prevention, as a deeper understanding of circumstances can contribute to more effective interventions.
Who is affected?
NLP and public health researchers, AI developers, suicide researchers, and data scientists working with text analysis and machine learning. Those developing and using LLMs for information extraction are directly affected by the study's insights on prompting and model selection.
What else you should know
The work focuses on datasets from the US National Violent Death Reporting System (NVDRS), a system that collects detailed information on violent deaths.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka skillnader observerades mellan modelltyperna?
Påverkar det EU?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Study compares performance of LLMs and fine-tuned models in "