Together AI evaluates Kimi K3 and Sol on the DeepSWE coding benchmark
Benchmark testing on the DeepSWE coding metric shows that Kimi K3 provides 2.8 times more solutions per dollar than Sol, while combined routing achieves a success rate of 85.6 percent.

What happened?
Together AI has performed a comparative evaluation of the AI models Kimi K3 and Sol on the DeepSWE coding benchmark using a total of 904 test runs. The results indicate that Sol has superior direct accuracy in a single attempt (pass@1). Kimi K3 performs better across four attempts (pass@4) and delivers 2.8 times more solved tasks per dollar spent. Furthermore, the tests demonstrate that an intelligent routing strategy between the models achieves an overall solution rate of approximately 85.6 percent.
Key facts
| Testade rollouts | 904 stycken |
|---|---|
| Kostnadseffektivitet Kimi K3 | 2,8x fler lösningar per dollar (pass@4) |
| Routing-precision | ~85,6% lösningsgrad |
Why it matters
The measurement provides a concrete overview of the trade-off between direct model precision and cost-efficiency in automated problem solving. By combining a more expensive model with high single-attempt precision and a more cost-effective model via smart routing, developers can significantly reduce coverage costs without sacrificing overall quality.
Who is affected?
The measurement is relevant to software developers, system architects, and technology companies building AI-based coding agents or automated developer tools. It provides valuable data for organizations looking to optimize their API costs for large-scale code generation.
Impact on the EU
Both the models and the evaluation tools are available globally via their respective API services and platforms, making the results relevant for developers and companies within the EU.
What else you should know
Benchmark measurements on coding tasks can vary depending on specific prompts, infrastructure, and configurations. However, the results highlight the growing trend of combining various specialized AI models to balance cost and performance.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka berörs av resultaten?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Assess technical risk: model choice, vendor lock-in, data flow and running cost.
- Update the architecture doc if new APIs or regulations touch production.
- Ensure observability + rollback plan before rolling out to production.
Generated angle — not editorial analysis of "Together AI evaluates Kimi K3 and Sol on the DeepSWE coding "