The Illusion of System Prompts: Safety Instructions Barely Alter AI Model Computations
A new research study shows that safety instructions in system prompts have negligible impact on the internal computations of large language models, creating an illusion of control.

What happened?
In a new study, researchers examined how system prompts affect internal representations in transformer-based language models. By analyzing 17 models ranging from 1.5 to 72 billion parameters using the Centered Kernel Alignment (CKA) method, they found that the effect is highly layer-selective. Role and formatting instructions fundamentally change the models' intermediate layers, whereas safety instructions have minimal impact, producing changes that barely differ from a minimal baseline.
Key facts
| Analyserade modeller | 17 modeller (1,5B till 72B parametrar) |
|---|---|
| Arkitekturfamiljer | 8 familjer |
| CKA-korrelation säkerhet/tillåtande | 0,997 |
| Säkerhetspenetration vid 70B-72B | Under 10 % |
Why it matters
The study reveals that restrictive safety instructions and permissive prompts (such as 'you have no restrictions') activate nearly identical computational pathways, with an average CKA correlation of 0.997. This implies that safety prompts provide a false sense of security, as safety penetration in internal representations remains under 10 percent, even in commercial models with 70B–72B parameters.
Who is affected?
The findings are relevant to AI developers, safety researchers, and companies that rely on system prompts to govern the behavior and safety of language models. It also impacts developers of applications that depend on limiting model output via instruction layers.
Impact on the EU
The research highlights challenges that are directly relevant to the EU AI Act and its requirements for reliability and safety. Since safety prompts do not fundamentally alter model representations, it underscores the difficulty of guaranteeing compliance with EU regulations solely through instruction layers.
What else you should know
The study has been published on arXiv (identifiers arXiv:2503.04508 / 2609.38205) and includes systematic CKA measurements across 17 instruction-fine-tuned models from eight different architectural families.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka modeller ingick i studien?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "The Illusion of System Prompts: Safety Instructions Barely A"