New method for faster multimodal AI models without performance loss
Researchers have developed a new method to accelerate multimodal AI models, enabling sustained efficiency even in complex reasoning tasks without the usual computational overhead.

What happened?
A new research study presents a method to streamline multimodal AI models that previously relied on explicit Chain-of-Thought (CoT) reasoning. Instead of generating entire CoT sequences, latent "think tokens" are introduced to represent the reasoning implicitly. These tokens are optimised via a CoT generation loss, while subsequent embedded tokens are optimised through a contrastive loss, resulting in high-performance, reasoning-aware representations.
Key facts
| Publikationsdatum | 24 maj 2026 |
|---|---|
| Forskningsområde | Artificiell Intelligens (cs.AI) |
| DOI | 2605.16638v1 |
”Recent research has demonstrated that Universal Multimodal Embedding (UME) benefits significantly from Chain-of-Thought (CoT) reasoning.”
”Despite its effectiveness, the computational overhead of generating explicit CoT traces is often prohibitive.”
”By optimizing think tokens using CoT generation loss and subsequent embedding tokens using contrastive loss, we produce high-performance, reasoning-aware representations at a constant inference cost.”
Why it matters
Traditional CoT methods, where a generative model creates explicit reasoning paths, have proven effective for multimodal embeddings. However, these entail significant computational overhead. The new method aims to retain the benefits of reasoning-based AI but without the previous performance losses in the form of increased inference costs. This means complex AI applications could become more accessible and faster to execute.
Who is affected?
The method primarily impacts AI researchers and developers working with multimodal models and Universal Multimodal Embeddings (UME). Reduced computational costs could lead to more efficient development and deployment of AI systems for image, text, and other data types. In the long term, this facilitates the development of more robust and applicable AI systems. Companies using AI models for multimodal tasks may also benefit from handling complex models at a lower cost and with faster response times.
What else you should know
The study also examines two architectural design questions: how best to extract "think tokens" and embedding tokens. The extract does not provide information on exact performance gains, nor does it quantify the performance loss of traditional CoT beyond describing it as "prohibitive".
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vem påverkas av detta?
Påverkar det EU?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "New method for faster multimodal AI models without performan"