Skip to content
Kodning & Utveckling· NewsAvailable

Together AI evaluates Kimi K3 and Sol on the DeepSWE coding benchmark

Benchmark testing on the DeepSWE coding metric shows that Kimi K3 provides 2.8 times more solutions per dollar than Sol, while combined routing achieves a success rate of 85.6 percent.

By the Aheadline editorial team·30 juli 2026·2 min read·Source: Together AI BlogVerifierad signalAI-generated
Together AI evaluates Kimi K3 and Sol on the DeepSWE coding benchmark
Together AI evaluates Kimi K3 and Sol on the DeepSWE coding benchmark
Together AI evaluates Kimi K3 and Sol on the DeepSWE coding benchmark
By · Policy- & EU-reporter
Last updated

What happened?

Together AI has performed a comparative evaluation of the AI models Kimi K3 and Sol on the DeepSWE coding benchmark using a total of 904 test runs. The results indicate that Sol has superior direct accuracy in a single attempt (pass@1). Kimi K3 performs better across four attempts (pass@4) and delivers 2.8 times more solved tasks per dollar spent. Furthermore, the tests demonstrate that an intelligent routing strategy between the models achieves an overall solution rate of approximately 85.6 percent.

Key facts

Testade rollouts904 stycken
Kostnadseffektivitet Kimi K32,8x fler lösningar per dollar (pass@4)
Routing-precision~85,6% lösningsgrad

Why it matters

The measurement provides a concrete overview of the trade-off between direct model precision and cost-efficiency in automated problem solving. By combining a more expensive model with high single-attempt precision and a more cost-effective model via smart routing, developers can significantly reduce coverage costs without sacrificing overall quality.

Who is affected?

The measurement is relevant to software developers, system architects, and technology companies building AI-based coding agents or automated developer tools. It provides valuable data for organizations looking to optimize their API costs for large-scale code generation.

Impact on the EU

Both the models and the evaluation tools are available globally via their respective API services and platforms, making the results relevant for developers and companies within the EU.

What else you should know

Benchmark measurements on coding tasks can vary depending on specific prompts, infrastructure, and configurations. However, the results highlight the growing trend of combining various specialized AI models to balance cost and performance.

Frequently asked questions

Quick answers about this story

Vad har hänt?
Together AI har genomfört 904 utvärderingskörmoment på riktmärket DeepSWE för att jämföra prestanda och kostnadseffektivitet mellan AI-modellerna Kimi K3 och Sol.
När hände det?
Testresultaten publicerades av Together AI i maj 2024.
Varför spelar det roll?
Testerna visar att kostnadseffektiva modeller som Kimi K3 kan ge fler lösta uppgifter per dollar, och att dynamisk routing mellan modeller kan optimera både prestanda och budget vid automatiserad kodgenerering.
Vilka berörs av resultaten?
Resultaten är direkt relevanta för alla utvecklare och företag globalt som använder API-baserade AI-modeller för programvaruutveckling och kodgenerering.
Original source
Together AI Blog·together.ai

The link opens in a new window and leads to the publisher's own site.

Verifierad signal

Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.

AI-verktyg i artikeln

Topics

#GPT-5.6 Sol#AI-benchmarking#Kimi K3#Kodgenerering
[ STAY UP TO DATE ]

Get similar news straight to your inbox

No affiliate linksCancel anytimeGDPR-friendly
[ Frequency ]
[ What do you want to read about? ]

You'll receive updates on 2 topics.

The reader's room

Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.

Sign in to submit a comment or question.

Loading comments…
How this affects you

Read the article through your role

  • Assess technical risk: model choice, vendor lock-in, data flow and running cost.
  • Update the architecture doc if new APIs or regulations touch production.
  • Ensure observability + rollback plan before rolling out to production.

Generated angle — not editorial analysis of "Together AI evaluates Kimi K3 and Sol on the DeepSWE coding "