Skip to content
Forskning· Analysis

New fine-tuning method reduces catastrophic forgetting in LLMs

A new method, Sparse Memory Finetuning (SMF), shows promising results in reducing "catastrophic forgetting" during the fine-tuning of large language models, according to a study published on arXiv.

By the Aheadline editorial team·7 juli 2026·2 min read·Source: arXiv cs.CL (NLP/LLM)Verifierad signalAI-generated
New fine-tuning method reduces catastrophic forgetting in LLMs
New fine-tuning method reduces catastrophic forgetting in LLMs
By · Policy- & EU-reporter
Last updated

What happened?

Researchers have introduced Sparse Memory Finetuning (SMF), a technique aimed at counteracting the phenomenon of "catastrophic forgetting" in pre-trained language models. The study, published on arXiv (ID: 2605.03229v1), demonstrates that SMF adds key-value memory layers to the model. During each training step, only a limited number of memory rows most relevant to the current data batch are updated.

Key facts

MetodSparse Memory Finetuning (SMF)
Publiceringsdatum26 maj 2026
Förbättring (MedMCQA)2,5 procentenheter
Glömska (generell kunskap)Inom 1 poäng från basmodell

SMF improves MedMCQA by 2.5 percentage points while keeping both forgetting probes within roughly 1 point of the base model, whereas LoRA and full finetuning achieve larger gains but with clear drift on both.

Forskarna, Forskare · arXiv

Why it matters

Catastrophic forgetting is a significant problem in AI, where models lose previously acquired knowledge when adapted to new tasks. The SMF method addresses this by selectively updating memory, preserving the model's general capacity while it learns new, specific tasks. This could lead to more robust and versatile AI systems.

Who is affected?

Researchers and developers working on fine-tuning large language models are the primary stakeholders. Companies implementing LLMs in their products could also benefit from more stable models that retain broad knowledge.

What else you should know

The study compared SMF with established methods such as LoRA and full fine-tuning. SMF improved performance on a medical question-answering task by 2.5 percentage points while keeping general knowledge loss (WikiText perplexity and TriviaQA accuracy) within one point of the base model. LoRA and full fine-tuning achieved greater improvements on the new task but with a distinct loss of general knowledge.

Frequently asked questions

Quick answers about this story

Vad har hänt?
En ny finjusteringsmetod kallad Sparse Memory Finetuning (SMF) har presenterats. Den syftar till att minska katastrofal glömska i stora språkmodeller genom att selektivt uppdatera minneslager under träningen.
När hände det?
Studien publicerades den 26 maj 2026 på arXiv under identifieraren 2605.03229v1.
Varför spelar det roll?
Katastrofal glömska är ett stort problem där LLM förlorar tidigare kunskap vid anpassning till nya uppgifter. SMF kan leda till mer robusta AI-modeller som behåller sin breda generella kunskap även efter finjustering för specifika applikationer.
Original source
arXiv cs.CL (NLP/LLM)·arxiv.org

The link opens in a new window and leads to the publisher's own site.

Verifierad signal

Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.

AI-verktyg i artikeln

Topics

#Models
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Decide whether this affects strategy over 6–12 months or is just noise.
  • Discuss with leadership: do we own the right question or does ownership need to move?
  • Ask: what risk are we taking by NOT acting on this this quarter?

Generated angle — not editorial analysis of "New fine-tuning method reduces catastrophic forgetting in LL"