Llama 4 Maverick Launches on Google Cloud with MoE Architecture
Google Cloud has launched Llama 4 Maverick 17B-128E, the largest model in the Llama 4 series. The model utilises Mixture-of-Experts architecture and offers capabilities in coding, reasoning, and image analysis.

What happened?
Google Cloud has announced the release of Llama 4 Maverick 17B-128E, a new large language model (LLM). As the most capable model in the Llama 4 series, it employs a Mixture-of-Experts (MoE) architecture and 'early fusion' techniques. Llama 4 Maverick is available via Google Cloud’s Gemini Enterprise Agent Platform as a partner model.
Key facts
”Llama 4 Maverick 17B-128E is Llama 4's largest and most capable model. It uses the Mixture-of-Experts (MoE) architecture and early fusion to provide coding, reasoning, and image capabilities.”
”Launch stage: GA Release date: April 29, 2025”
Why it matters
The launch of Llama 4 Maverick expands the range of advanced AI models available to developers and enterprises on the Google Cloud platform. The MoE architecture enables more efficient handling of complex tasks in coding, logical reasoning, and image understanding, potentially leading to more sophisticated AI applications. The model's ability to process multiple data types (text, code, images) reduces the reliance on multiple specialised models.
Who is affected?
Developers, enterprises building AI applications, and end-users of AI tools are affected. Specifically, developers gain access to a more powerful model for integration into projects, while enterprises stand to benefit from improved performance and efficiency in AI-driven solutions requiring text, code, and image processing. The model is designed for 'pay-as-you-go' and 'provisioned throughput' usage.
Impact on the EU
Llama 4 Maverick 17B-128E is initially available only in the US region (us-east5) and for ML processing (Multi-region) within the United States. This means EU-based users cannot currently deploy the model with low latency within the EU, which may impact data sovereignty and regulatory compliance for certain European enterprises.
What else you should know
The model features a context length of 524,288 tokens and a maximum output of 8,192 tokens. It supports features such as batch predictions, function calling, and structured output, though it currently lacks support for Llama Guard.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Påverkar det EU?
The link opens in a new window and leads to the publisher's own site.
Källan är en aggregator eller syndikering — vi rekommenderar att verifiera hos primärutgivaren.
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Llama 4 Maverick Launches on Google Cloud with MoE Architect"