Skip to content
Forskning· News

How Transformer models calculate responses in hidden layers

New research shows that Transformer models deliberately perform their calculations perpendicular to the readout axis to protect complex reasoning from vocabulary interference.

By the Aheadline editorial team·12 aug. 2026·2 min read·Source: arXiv cs.CL (NLP/LLM)Verifierad signalAI-generated
How Transformer models calculate responses in hidden layers
How Transformer models calculate responses in hidden layers
How Transformer models calculate responses in hidden layers
By · Policy- & EU-reporter
Last updated

What happened?

In a new study, researchers have demonstrated how Transformer-based AI models handle internal calculations. Instead of directly processing information in the same direction as the final output is read, early calculations occur at angles near 90 degrees from the readout axis. It is only in the final phase of the model that the answer is shifted to the main readout axis by adding information rather than rotating existing content.

Key facts

Vinkel för tidiga beräkningar75 till 96 grader från avläsningsaxeln
Skadefaktor vid felaktig placering64 till 84 gånger större än slumpmässig rotation
Modellarkitektur i studien12-lagers Transformer-modell

Why it matters

Previously, researchers viewed the angled internal representations of these models as an obstacle to interpreting what happens during computation. The new study shows that this separation serves a crucial function: it isolates and protects ongoing calculations from being disrupted by vocabulary representations. Forcing attention mechanisms to align with the readout axis significantly damages the model's ability to perform complex reasoning.

Who is affected?

The discovery is particularly significant for AI researchers, language model developers, and experts in mechanistic interpretability. The results provide a better foundation for understanding how AI models make decisions and how training methods can be optimized without compromising the models' capacity for complex logic.

Impact on the EU

The research concerns fundamental model architecture and the theoretical understanding of AI, affecting development and safety analysis globally. As this is basic research, there are no specific EU restrictions or market barriers directly linked to the discovery.

What else you should know

The researchers also show that when all layers are forced to align with the final output—often done in techniques such as early exit—the model loses the ability to perform multi-step reasoning. Although simpler tests, such as language modeling or word prediction, appear unaffected, performance on complex tasks collapses. This underscores that geometric separation in vector space is essential for deeper inferential capability.

Frequently asked questions

Quick answers about this story

Vad har hänt?
Forskare har upptäckt att Transformer-modeller utför sina beräkningar i vinklar som är nästan vinkelräta mot avläsningsaxeln, vilket skyddar sammanställningen av information från vokabuläret innan det slutgiltiga svaret genereras.
När hände det?
Studien publicerades i det öppna arkivet arXiv den 11 augusti 2026.
Varför spelar det roll?
Det förklarar hur språkmodeller faktiskt bearbetar information i sina dolda lager och visar varför vissa optimeringstekniker kan förstöra modellens förmåga till komplexa resonemang.
Vilka berörs av upptäckten?
Resultaten är främst relevanta för AI-forskare, utvecklare av stora språkmodeller och experter som arbetar med tolkningsbarhet och säkerhet inom AI.
Original source
arXiv cs.CL (NLP/LLM)·arxiv.org

The link opens in a new window and leads to the publisher's own site.

Verifierad signal

Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.

AI-verktyg i artikeln

Topics

#AI-forskning#Large Language Models (LLMs)#Natural Language Processing (NLP)
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Which processes can be simplified or automated based on this?
  • Who trains the team — and when? Set a clear owner and deadline.
  • Follow up KPIs on lead time, quality and cost after adoption.

Generated angle — not editorial analysis of "How Transformer models calculate responses in hidden layers"