Skip to content
Chatt & Assistenter· NewsAvailable

Meta AI handles millions of queries with new efficient Llama architecture

Meta AI processes millions of queries daily using a 400-billion parameter model that activates only 17 billion per token.

By the Aheadline editorial team·1 aug. 2026·2 min read·Source: Entity-watch: Meta AIVerifierad signalAI-generated
Meta AI handles millions of queries with new efficient Llama architecture
Meta AI handles millions of queries with new efficient Llama architecture
Meta AI handles millions of queries with new efficient Llama architecture
By · Policy- & EU-reporter
Last updated

What happened?

Meta AI handles millions of user queries daily via existing applications within the Meta ecosystem. The underlying engine is based on a model with a total of 400 billion parameters. Through a Mixture-of-Experts architecture, each individual token is routed through only 17 billion active parameters at a time, representing a new approach for the Llama family.

Key facts

Totalt antal parametrar400 miljarder
Aktiva parametrar per token17 miljarder
Publiceringsdatum31 juli 2026

Why it matters

The method of activating only a small subset of parameters per token makes it possible to run a massive, high-capacity model at a fraction of standard computational costs. This marks a technical shift for the Llama series, which previously relied on fully activated dense models.

Who is affected?

Developers, AI researchers, and end-users are affected by the increased computational efficiency of large-scale AI models. The change enables Meta to offer faster responses to millions of daily queries in apps such as Instagram, WhatsApp, and Facebook.

Impact on the EU

Llama models are distributed globally as open source, but large-scale AI services and data processing in Meta apps are strictly regulated in Europe under the EU AI Act and GDPR. Local adjustments may therefore affect the rollout of specific Meta AI features in the EU.

What else you should know

MoE architectures have become the industry standard for scaling language models without runaway computational costs per generated token. By activating only 17 billion parameters per token, Meta achieves significantly lower latency and energy consumption when handling immense query volumes.

Frequently asked questions

Quick answers about this story

Vad har hänt?
Meta har integrerat en AI-modell på 400 miljarder parametrar i sina appar, vilken använder en Mixture-of-Experts-arkitektur för att aktivera endast 17 miljarder parametrar per token.
När hände det?
Meta rapporterade om den nya Llama-arkitekturen och dess kapacitet den 31 juli 2026.
Varför spelar det roll?
Genom att enbart aktivera 17 miljarder parametrar åt gången kan Meta köra en avancerad 400-miljardersmodell för miljontals användare med avsevärt lägre beräkningsresurser och högre hastighet.
Vilka berörs av teknikskiftet?
Modellens effektiva arkitektur minskar beräkningskostnaderna för utvecklare och levererar snabbare svar till miljontals användare i Metas appar.
Original source
Entity-watch: Meta AI·gcn.com

The link opens in a new window and leads to the publisher's own site.

Verifierad signal

Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.

AI-verktyg i artikeln

Topics

#Large Language Models (LLMs)#AI-assistenter#Meta AI
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Decide whether this affects strategy over 6–12 months or is just noise.
  • Discuss with leadership: do we own the right question or does ownership need to move?
  • Ask: what risk are we taking by NOT acting on this this quarter?

Generated angle — not editorial analysis of "Meta AI handles millions of queries with new efficient Llama"