Meta AI handles millions of queries with new efficient Llama architecture
Meta AI processes millions of queries daily using a 400-billion parameter model that activates only 17 billion per token.

What happened?
Meta AI handles millions of user queries daily via existing applications within the Meta ecosystem. The underlying engine is based on a model with a total of 400 billion parameters. Through a Mixture-of-Experts architecture, each individual token is routed through only 17 billion active parameters at a time, representing a new approach for the Llama family.
Key facts
| Totalt antal parametrar | 400 miljarder |
|---|---|
| Aktiva parametrar per token | 17 miljarder |
| Publiceringsdatum | 31 juli 2026 |
Why it matters
The method of activating only a small subset of parameters per token makes it possible to run a massive, high-capacity model at a fraction of standard computational costs. This marks a technical shift for the Llama series, which previously relied on fully activated dense models.
Who is affected?
Developers, AI researchers, and end-users are affected by the increased computational efficiency of large-scale AI models. The change enables Meta to offer faster responses to millions of daily queries in apps such as Instagram, WhatsApp, and Facebook.
Impact on the EU
Llama models are distributed globally as open source, but large-scale AI services and data processing in Meta apps are strictly regulated in Europe under the EU AI Act and GDPR. Local adjustments may therefore affect the rollout of specific Meta AI features in the EU.
What else you should know
MoE architectures have become the industry standard for scaling language models without runaway computational costs per generated token. By activating only 17 billion parameters per token, Meta achieves significantly lower latency and energy consumption when handling immense query volumes.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka berörs av teknikskiftet?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Meta AI handles millions of queries with new efficient Llama"