Moonshot AI Releases Kimi K3: Running at Over 20 TPS on GB10 Clusters
Moonshot AI has released the weights for its open-weight language model, Kimi K3. Meanwhile, developers report the model is achieving speeds of over 20 tokens per second on GB10 clusters.

What happened?
Chinese company Moonshot AI has published the model weights for its language model, Kimi K3, establishing it as one of the world's largest open-weight AI models. According to reports from the developer community, the full model has been successfully run on a cluster of 16 Nvidia GB10 units with speeds exceeding 20 tokens per second. Shortly after the release, Moonshot AI was forced to pause new registrations to its proprietary service due to intense user demand.
Key facts
| Modell | Moonshot AI Kimi K3 |
|---|---|
| Status vikter | Öppet tillgängliga (open-weight) |
| Hårdvarukluster | 16x GB10 |
| Prestanda | 20+ tokens per sekund |
Why it matters
The release of Kimi K3 marks a significant milestone for open AI models, offering performance that challenges leading closed-source commercial alternatives. However, the launch triggered a massive influx of users, resulting in capacity issues for Moonshot AI and highlighting the substantial infrastructure requirements needed to operate the model at scale.
Who is affected?
The release primarily impacts AI researchers, system architects, and developers building applications based on large language models. Companies seeking to run proprietary AI models without transmitting data to third-party clouds are also affected, as Kimi K3 provides access to advanced performance on internal infrastructure.
Impact on the EU
Because Moonshot AI has published Kimi K3 with open weights, developers and companies within the EU can host the model in local data centres. This facilitates compliance with GDPR and the forthcoming EU AI Act regarding data localisation and privacy.
What else you should know
Reports within the community indicate that running the full model requires significant hardware infrastructure due to its scale. Utilising a cluster of 16 GB10 units to achieve over 20 tokens per second underscores the extreme computational demands required to run next-generation open-weight language models locally.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Påverkar det EU?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Assess technical risk: model choice, vendor lock-in, data flow and running cost.
- Update the architecture doc if new APIs or regulations touch production.
- Ensure observability + rollback plan before rolling out to production.
Generated angle — not editorial analysis of "Moonshot AI Releases Kimi K3: Running at Over 20 TPS on GB10"