OpenAI Launches New Realtime API Models for Voice Agents
OpenAI has released gpt-realtime-2.1 and gpt-realtime-2.1-mini, reducing latency in the Realtime API for voice agents by at least 25% and enhancing interaction quality.

What happened?
OpenAI recently launched two new models, gpt-realtime-2.1 and gpt-realtime-2.1-mini, for its Realtime API. These updates aim to improve the performance of AI-powered voice agents. Latency on the platform has been reduced by at least 25% for the 95th percentile across all existing Realtime voice models through improved prompt caching. The new mini model also introduces a distilled reasoning model capable of maintaining conversation during tool calls, effectively eliminating silent periods.
Key facts
| Nya modeller | gpt-realtime-2.1 och gpt-realtime-2.1-mini |
|---|---|
| Latency-reduktion (p95) | Minst 25% |
| Lanseringsdatum | 7 juli 2026 |
”OpenAI released two new models to its Realtime API on Sunday night, cutting a latency problem that had made AI phone agents feel broken even when they were working correctly.”
”Platform-wide, OpenAI said improved prompt caching reduced p95 latency by at least 25% across all Realtime voice models — meaning the slowest 5% of response times, the tail that defines whether a voice system feels reliable, became materially faster for every existing integration”
Why it matters
Reduced latency is critical for voice agents to be perceived as natural and efficient. Previous issues with delays have made AI voice agents appear inadequate. By eliminating long silences during computation and improving the handling of audio and interruptions, interactions become more fluid and user-friendly. This is a significant development for increasing the acceptance and functionality of AI-based voice systems.
Who is affected?
Developers building AI-driven voice agents are directly affected, as they gain access to more responsive and capable tools. Companies implementing customer service or other voice-controlled systems will offer a better user experience. End-users interacting with AI voice agents will experience smoother and more natural conversations.
What else you should know
The gpt-realtime-2.1 model improves the handling of numbers, background noise, and mid-sentence interruptions. This architecture, utilizing a single model for both speech processing and synthesis, differs from traditional systems that often rely on several separate services.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka bolag berörs?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "OpenAI Launches New Realtime API Models for Voice Agents"