Skip to content
Forskning· Analysis

The Illusion of System Prompts: Safety Instructions Barely Alter AI Model Computations

A new research study shows that safety instructions in system prompts have negligible impact on the internal computations of large language models, creating an illusion of control.

By the Aheadline editorial team·1 okt. 2026·2 min read·Source: arXiv cs.CL (NLP/LLM)Verifierad signalAI-generated
The Illusion of System Prompts: Safety Instructions Barely Alter AI Model Computations
The Illusion of System Prompts: Safety Instructions Barely Alter AI Model Computations
The Illusion of System Prompts: Safety Instructions Barely Alter AI Model Computations
By · Policy- & EU-reporter
Last updated
Vad betyder det för mig?

What happened?

In a new study, researchers examined how system prompts affect internal representations in transformer-based language models. By analyzing 17 models ranging from 1.5 to 72 billion parameters using the Centered Kernel Alignment (CKA) method, they found that the effect is highly layer-selective. Role and formatting instructions fundamentally change the models' intermediate layers, whereas safety instructions have minimal impact, producing changes that barely differ from a minimal baseline.

Key facts

Analyserade modeller17 modeller (1,5B till 72B parametrar)
Arkitekturfamiljer8 familjer
CKA-korrelation säkerhet/tillåtande0,997
Säkerhetspenetration vid 70B-72BUnder 10 %

Why it matters

The study reveals that restrictive safety instructions and permissive prompts (such as 'you have no restrictions') activate nearly identical computational pathways, with an average CKA correlation of 0.997. This implies that safety prompts provide a false sense of security, as safety penetration in internal representations remains under 10 percent, even in commercial models with 70B–72B parameters.

Who is affected?

The findings are relevant to AI developers, safety researchers, and companies that rely on system prompts to govern the behavior and safety of language models. It also impacts developers of applications that depend on limiting model output via instruction layers.

Impact on the EU

The research highlights challenges that are directly relevant to the EU AI Act and its requirements for reliability and safety. Since safety prompts do not fundamentally alter model representations, it underscores the difficulty of guaranteeing compliance with EU regulations solely through instruction layers.

What else you should know

The study has been published on arXiv (identifiers arXiv:2503.04508 / 2609.38205) and includes systematic CKA measurements across 17 instruction-fine-tuned models from eight different architectural families.

Frequently asked questions

Quick answers about this story

Vad har hänt?
Forskare publicerade en ny analys av hur systemprompter påverkar språkmodellers interna beräkningar via Centered Kernel Alignment (CKA).
När hände det?
Studien publicerades som ett preprint på arXiv i mars 2025.
Varför spelar det roll?
Det visar att säkerhetsinstruktioner i systemprompter har minimal effekt på modellens djupare representationer, vilket innebär att prompter inte räcker som säkerhetsspärr.
Vilka modeller ingick i studien?
Analysen omfattade 17 instruktionsfinjusterade modeller från 8 arkitekturfamiljer, i storlekar från 1,5B till 72B parametrar.
Original source
arXiv cs.CL (NLP/LLM)·arxiv.org

The link opens in a new window and leads to the publisher's own site.

Verifierad signal

Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.

AI-verktyg i artikeln

Topics

#Safety#Models
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Decide whether this affects strategy over 6–12 months or is just noise.
  • Discuss with leadership: do we own the right question or does ownership need to move?
  • Ask: what risk are we taking by NOT acting on this this quarter?

Generated angle — not editorial analysis of "The Illusion of System Prompts: Safety Instructions Barely A"