Models

OpenAI ships three realtime audio models, led by GPT-Realtime-2

OpenAI launches GPT-Realtime-2, GPT-Realtime-Translate, and GPT-Realtime-Whisper in the Realtime API, bringing reasoning, live translation, and streaming transcription to voice apps.

Advancing voice intelligence with new models in the API
Advancing voice intelligence with new models in the APIGauravonomics / Openverse
By James Calloway6 min read

Updated

Why it matters

  • GPT-Realtime-2 scores 15.2% higher on Big Bench Audio than GPT-Realtime-1.5 and expands the context window from 32K to 128K tokens.
  • GPT-Realtime-Translate supports 70+ input languages and 13 output languages at $0.034 per minute; GPT-Realtime-Whisper costs $0.017 per minute.
  • Zillow reports a 26-point lift in call success rate (95% vs. 69%) on its hardest adversarial benchmark with GPT-Realtime-2.

OpenAI has launched three audio models in its Realtime API: GPT‑Realtime‑2, its first voice model with GPT‑5‑class reasoning; GPT‑Realtime‑Translate, which translates speech from more than 70 input languages into 13 output languages while keeping pace with the speaker; and GPT‑Realtime‑Whisper, a streaming speech-to-text model that transcribes audio live as people talk.

The release matters because voice is becoming a primary way people use software — asking for help while driving, changing a travel plan in an airport, or getting support in a preferred language — and OpenAI is betting that useful voice products need more than fast turn-taking. A voice agent, the company argues, must understand intent, track context, recover when a request changes, and use tools while the conversation continues. Together, the three models move realtime audio from simple call-and-response toward voice interfaces that can listen, reason, translate, transcribe, and act as a conversation unfolds.

Three patterns, three named customers

OpenAI says developers are building around three emerging patterns in voice AI.

Voice-to-action lets people describe a need and have the system reason through it, use tools, and complete the task. Zillow is building an assistant that handles requests like "find me homes within my BuyAbility, avoid busy streets, and schedule a tour for Saturday."

Systems-to-voice turns software context into live spoken guidance. A travel app could tell a traveler: "Your inbound flight is delayed, but you can still make your connection. I found the new gate, mapped the fastest route through the terminal, and your bag is still expected to transfer."

Voice-to-voice keeps live conversations going across languages and shifting context. Deutsche Telekom is building voice support where customers speak in the language they're most comfortable with while the model translates in real time.

The patterns can combine. Priceline is working toward travelers managing entire trips by voice — searching for flights and hotels conversationally, adjusting a hotel reservation after a delay, checking TSA wait times, and translating conversations on the ground.

GPT-Realtime-2: reasoning, tools, and a 128K context window

GPT‑Realtime‑2 is built for live interactions where the model keeps the conversation moving while it reasons through a request, calls tools, handles corrections and interruptions, and adjusts its delivery. Several capabilities stand out.

Developers can enable preambles — short phrases like "let me check that" or "one moment while I look into it" — so users know the agent is working. The model supports parallel tool calls with tool transparency, announcing actions audibly with phrases like "checking your calendar" or "looking that up now." It shows stronger recovery behavior, saying "I'm having trouble with that right now" instead of failing silently.

The context window grows from 32K to 128K, supporting longer, more coherent sessions and more complex task flows. The model better retains specialized terminology, proper nouns, and healthcare terms. Its tone is more controllable — calm while resolving an issue, empathetic when a user is frustrated, upbeat when confirming a success. Developers can select reasoning effort from minimal, low, medium, high, and xhigh levels, with low as the default, balancing latency against deliberation.

The benchmarks back the claims. GPT‑Realtime‑2 at high effort scores 15.2% higher on Big Bench Audio for audio intelligence than GPT‑Realtime‑1.5. At xhigh, it scores 13.8% higher on Audio MultiChallenge for instruction following, with stronger reasoning, context management, and control in live conversations.

Zillow put the model through adversarial testing. "What stood out about GPT-Realtime-2 was the intelligence and tool-calling reliability it brings to complex voice interactions. On our hardest adversarial benchmark, this translates to a 26-point lift in call success rate after prompt optimization (95% vs. 69%). GPT-Realtime-2 is also materially more robust on Fair Housing compliance, which is critical for our business. The combination of agentic competence and guardrail strength is what makes it viable for production voice at Zillow," said Josh Weisberg, SVP and Head of AI at Zillow.

GPT-Realtime-Translate: 70+ input languages, 13 outputs

Live translation must preserve meaning while keeping pace with speakers who talk naturally, switch context, use regional pronunciation, and lean on domain-specific language. GPT‑Realtime‑Translate targets exactly that, letting each person speak in a preferred language, hear the conversation translated in real time, and read live transcriptions. OpenAI positions it for customer support, cross-border sales, education, events, media, and creator platforms with global audiences.

Deutsche Telekom is testing the model for multilingual voice interactions, where lower latency and stronger fluency can make cross-language conversations feel more natural. Vimeo demonstrated the model translating a product education video live as it plays, sparing global customers a separately produced version.

BolnaAI, which builds voice AI for India, ran evaluations across Hindi, Tamil, and Telugu. "In our evals across Hindi, Tamil, and Telugu, GPT-Realtime-Translate delivered 12.5% lower Word Error Rates than any other model we tested, along with lower fallback rates, higher task completion, and latency that sustained natural conversation. It sets a new standard for multilingual voice AI," said Prateek Sachan, Co-founder & CTO at BolnaAI.

GPT-Realtime-Whisper: speech-to-text as it happens

GPT‑Realtime‑Whisper transcribes audio while people speak, aimed at live captions, meeting notes that keep up with the conversation, voice agents that must understand users continuously, and follow-up workflows in customer support, healthcare, sales, and recruiting — anywhere high-volume spoken interactions need to become usable data in real time.

Safety and compliance

The Realtime API runs active classifiers over sessions, meaning certain conversations can be halted if they violate OpenAI's harmful content guidelines. Developers can add their own guardrails through the Agents SDK. Usage policies prohibit repurposing outputs for spam or deception, and developers must disclose to end users that they're interacting with AI unless it's obvious from context. The API fully supports EU Data Residency for EU-based applications and is covered by OpenAI's enterprise privacy commitments — a relevant detail for European deployments at a time when data sovereignty rules increasingly shape AI procurement.

Pricing and availability

All three models are available now in the Realtime API. GPT‑Realtime‑2 costs $32 per 1M audio input tokens ($0.40 for cached input tokens) and $64 per 1M audio output tokens. GPT‑Realtime‑Translate is priced at $0.034 per minute, and GPT‑Realtime‑Whisper at $0.017 per minute. Developers can test the models in the Playground or start building via a preloaded prompt in Codex.

With named customers already reporting production-grade results — Zillow's 95% call success rate on its hardest benchmark — OpenAI is signaling that realtime voice is ready to move from demo to deployment, and the token and per-minute pricing gives competitors in voice AI a concrete bar to beat.

Original: cdn.openai.com

Share this article:

More from James Calloway

James Calloway

Show full bio

News editor covering industry trends and analytics at AI In Context.

121 articles

Related articles

  1. OpenAI Launches gpt-realtime Speech-to-Speech Model With MCP and Phone Support
  2. OpenAI puts GPT-Live-1 in the API at $0.05 per minute
  3. OpenAI Kills the Turn Detector: Inside GPT-Live's Realtime Voice Architecture
  4. OpenAI Releases GPT-5.2, Its New Frontier Model for Professional Work
  5. Google's Flash TTS models build AI voices from text descriptions

« Previous articleNext article »