New AI method locks safety circuits during self-improving evolution
New research introduces Circuit-Anchored Evolution, a method inspired by biological evolution that secures the safety circuits of AI models during self-improving training.

What happened?
Researchers have introduced a method called Circuit-Anchored Evolution (CAE) to prevent AI models from losing their safety guardrails during self-improving evolutionary processes. By using mechanistic interpretability, a safety circuit comprising less than two percent of the model's features is identified. This specific component is locked during optimisation while the remaining parts of the model are permitted to adapt and develop new capabilities.
Key facts
| Säkerhetskretsens storlek | Mindre än 2% av modellens funktioner |
|---|---|
| Metodens namn | Circuit-Anchored Evolution (CAE) |
| Publiceringsdatum | 10 augusti 2026 |
Why it matters
When language models are trained through self-evolution, there is a risk that optimisation will prioritise capability alone, leading to the degradation of safety functions. By drawing inspiration from biological evolution and so-called developmental constraints, the research demonstrates how core functions can be mathematically anchored. This enables continuous improvement of a model's capacity without intentionally or unintentionally manipulating its fundamental safety structures.
Who is affected?
The method concerns AI researchers and developers working with self-improving language models and automatic model optimisation. It is particularly relevant for entities developing advanced systems where safety mechanisms need to be automatically maintained over time.
Impact on the EU
The method addresses fundamental safety mechanisms in large language models and is directly relevant to requirements for risk management and AI safety under the EU AI Act, although the research paper itself does not address specific EU legislation.
What else you should know
The method relies on mechanistic interpretability to isolate and protect the specific features within the model's architecture that govern safety behaviours. The researchers draw a direct parallel to developmental biology, where preserved gene structures allow for the adaptation of organisms without the loss of essential life-sustaining functions.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Hur kopplar metoden till biologisk evolution?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "New AI method locks safety circuits during self-improving ev"