Meta Launches Llama 4: MoE Architecture and Record-Breaking Context Window
Meta has launched the Llama 4 family featuring MoE architecture and native multimodality. The models support context windows of up to 10 million tokens and demonstrate high performance on local Apple hardware.
What happened?
Meta has launched the Llama 4 model family, marking the company’s transition to MoE (Mixture of Experts) architecture and native multimodality. The family consists of three models: Llama 4 Scout (17B active parameters, 16 experts, 109B total), Llama 4 Maverick (17B active parameters, 128 experts, 402B total), and Behemoth (288B active parameters, 16 experts, 2T total). The models support a context window of up to 10 million tokens.
Key facts
| Lanseringsdatum | 7 april 2025 |
|---|---|
| Maximalt kontextfönster | 10 miljoner tokens |
| Största modellens parametrar | 2 000 miljarder (2T) totalt |
| Molnstöd | Databricks Foundation Model API |
Why it matters
The launch represents a significant leap for open-source AI by combining an extremely long context window with MoE architecture. By activating only a fraction of the parameters during each inference run, computational requirements are reduced, making it possible to run extremely large models on consumer hardware and local infrastructure.
Who is affected?
Developers, AI researchers, and companies wishing to run large-scale open-source models locally are directly affected. The efficient MoE architecture enables the execution of very large models on local hardware, such as Mac computers with Apple Silicon and high unified memory capacity.
Impact on the EU
The Llama 4 models are released as open source and are available globally, including within the EU. However, full-scale deployment in commercial services within the EU may be subject to future compliance requirements under the EU AI Act.
What else you should know
Initial tests showed that Llama 4 Maverick achieved a generation speed of 50 tokens per second on a single Mac equipped with an M3 Ultra chip. Conversely, independent tests showed varying results regarding code generation, with some users reporting shortcomings in complex coding tasks compared to specialized models.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Påverkar det EU?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Meta Launches Llama 4: MoE Architecture and Record-Breaking "