Skip to content
Kodning & Utveckling· NewsAvailable

Moonshot AI Releases Kimi K3: Running at Over 20 TPS on GB10 Clusters

Moonshot AI has released the weights for its open-weight language model, Kimi K3. Meanwhile, developers report the model is achieving speeds of over 20 tokens per second on GB10 clusters.

By the Aheadline editorial team·5 aug. 2026·2 min read·Source: Reddit r/LocalLLaMAVerifierad signalAI-generated
Moonshot AI Releases Kimi K3: Running at Over 20 TPS on GB10 Clusters
Moonshot AI Releases Kimi K3: Running at Over 20 TPS on GB10 Clusters
By · Policy- & EU-reporter

What happened?

Chinese company Moonshot AI has published the model weights for its language model, Kimi K3, establishing it as one of the world's largest open-weight AI models. According to reports from the developer community, the full model has been successfully run on a cluster of 16 Nvidia GB10 units with speeds exceeding 20 tokens per second. Shortly after the release, Moonshot AI was forced to pause new registrations to its proprietary service due to intense user demand.

Key facts

ModellMoonshot AI Kimi K3
Status vikterÖppet tillgängliga (open-weight)
Hårdvarukluster16x GB10
Prestanda20+ tokens per sekund

Why it matters

The release of Kimi K3 marks a significant milestone for open AI models, offering performance that challenges leading closed-source commercial alternatives. However, the launch triggered a massive influx of users, resulting in capacity issues for Moonshot AI and highlighting the substantial infrastructure requirements needed to operate the model at scale.

Who is affected?

The release primarily impacts AI researchers, system architects, and developers building applications based on large language models. Companies seeking to run proprietary AI models without transmitting data to third-party clouds are also affected, as Kimi K3 provides access to advanced performance on internal infrastructure.

Impact on the EU

Because Moonshot AI has published Kimi K3 with open weights, developers and companies within the EU can host the model in local data centres. This facilitates compliance with GDPR and the forthcoming EU AI Act regarding data localisation and privacy.

What else you should know

Reports within the community indicate that running the full model requires significant hardware infrastructure due to its scale. Utilising a cluster of 16 GB10 units to achieve over 20 tokens per second underscores the extreme computational demands required to run next-generation open-weight language models locally.

Frequently asked questions

Quick answers about this story

Vad har hänt?
Moonshot AI släppte modellvikterna för sin öppna språkmodell Kimi K3, samtidigt som utvecklare rapporterat att den fulla modellen kunnat köras i över 20 tokens per sekund på ett kluster med 16 Nvidia GB10-enheter.
När hände det?
Modellvikterna släpptes och rapporterades i utvecklarforum i början av 2026, samtidigt som det höga trycket tvingade Moonshot AI att pausa nya användarregistreringar.
Varför spelar det roll?
Det spelar roll eftersom Kimi K3 utgör en av världens största öppet tillgängliga modeller, vilket ger utvecklare och företag möjlighet att köra avancerad AI på egen hårdvara utan beroende av stängda molntjänster.
Påverkar det EU?
Ja, eftersom vikterna är öppna kan organisationer i EU fritt ladda ned och köra modellen på egen infrastruktur i enlighet med europeiska dataskyddslagar.
Original source
Reddit r/LocalLLaMA·reddit.com

The link opens in a new window and leads to the publisher's own site.

Verifierad signal

Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.

AI-verktyg i artikeln

Topics

#GPU#Kimi K3#Large Language Models (LLMs)#AI-infrastruktur#AI-modell
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Assess technical risk: model choice, vendor lock-in, data flow and running cost.
  • Update the architecture doc if new APIs or regulations touch production.
  • Ensure observability + rollback plan before rolling out to production.

Generated angle — not editorial analysis of "Moonshot AI Releases Kimi K3: Running at Over 20 TPS on GB10"