Moonshot AI Releases Kimi K3: Running at Over 20 TPS on GB10 Clusters
Moonshot AI has released the weights for its open-language model, Kimi K3. Meanwhile, developers report the model is achieving speeds of over 20 tokens per second on GB10 clusters.

What happened?
Chinese-based Moonshot AI has published the model weights for its language model Kimi K3, making it one of the world's largest publicly available AI models. According to reports from the developer community, the full model has been successfully deployed on a cluster of 16 Nvidia GB10 units, achieving speeds exceeding 20 tokens per second. Shortly after the launch, Moonshot AI was forced to temporarily pause new registrations for its own service due to overwhelming demand.
Key facts
| Modell | Moonshot AI Kimi K3 |
|---|---|
| Status vikter | Öppet tillgängliga (open-weight) |
| Hårdvarukluster | 16x GB10 |
| Prestanda | 20+ tokens per sekund |
Why it matters
The release of Kimi K3 represents a significant milestone for open AI models, offering performance that challenges leading closed commercial models. However, the launch triggered a massive influx of users, resulting in capacity issues for Moonshot AI and highlighting the significant infrastructure requirements needed to run the model at scale.
Who is affected?
This release is of primary interest to AI researchers, system architects, and developers building applications based on large-scale language models. Companies seeking to run their own AI models without transmitting data to third-party clouds are also affected, as Kimi K3 provides access to advanced performance on private infrastructure.
Impact on the EU
Because Moonshot AI has published Kimi K3 as an open-weight model, developers and companies in the EU can host the model in local data centres. This facilitates compliance with GDPR and the upcoming EU AI Act regarding data localisation and privacy.
What else you should know
Community reports indicate that running the full model requires substantial hardware infrastructure due to its scale. Utilizing a cluster of 16 GB10 units to achieve over 20 tokens per second underscores the extreme computational requirements necessary to run next-generation open language models locally.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Påverkar det EU?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Assess technical risk: model choice, vendor lock-in, data flow and running cost.
- Update the architecture doc if new APIs or regulations touch production.
- Ensure observability + rollback plan before rolling out to production.
Generated angle — not editorial analysis of "Moonshot AI Releases Kimi K3: Running at Over 20 TPS on GB10"