Akashic presents MemAttention for more efficient LLM inference
A new memory architecture, Akashic with MemAttention, can significantly increase the efficiency of large language model inference by handling long contexts more dynamically.

What happened?
Researchers have introduced Akashic, a new memory system for large language models (LLMs), centred around the MemAttention technique. This system organises context into bounded "chunks" and models semantic relationships between them, eliminating the need to continuously rewrite the entire history. Akashic also implements a hardware-software co-designed memory placement to locate related data, which reduces fragmentation and I/O overhead.
Key facts
| Teknik | MemAttention |
|---|---|
| Förbättring i noggrannhet | Upp till 10,2 poäng |
| Ökad genomströmning | Upp till 1,21x |
”Recent LLM-based agent systems continuously accumulate context across multi-turn interactions, tool invocations, and cross-session workflows. Replaying the full history for every request quickly becomes impractical: long contexts increase prefill cost, may exceed context limits,”
Why it matters
The problem of long contexts in LLMs has been a significant challenge, as they increase "prefill" costs, risk exceeding context limits, and can degrade both efficiency and quality by tracking irrelevant information. Akashic addresses these shortcomings by selectively managing context memory, leading to improved task accuracy and throughput. The technology represents a step towards more scalable and cost-effective management of complex AI interactions.
Who is affected?
Primarily affected are developers and operators of LLM-based agent systems, particularly those working with applications requiring long and complex interactions. End-users of AI assistants and agents benefit indirectly through faster and more relevant responses, as underlying systems become more efficient and capable of handling broader dialogue contexts without performance degradation.
What else you should know
The presented research indicates up to a 10.2 point improvement in task accuracy and up to 1.21x higher throughput across four representative workloads and three model sizes.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka bolag berörs?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Akashic presents MemAttention for more efficient LLM inferen"