OpenAI puts GPT-Live-1 in the API at $0.05 per minute
OpenAI's GPT-Live-1 arrives in the API at $0.05 per minute, beating GPT-Realtime-2.1 by 30 points on Full Duplex Bench and cutting learner interruptions by nearly 80% at Speak.

Updated
Why it matters
- GPT-Live-1 is available in the API today at $0.05 per minute for the front-end voice layer, with developers choosing their own backend reasoning model.
- OpenAI's evaluations show GPT-Live-1 improving Full Duplex Bench performance by 30 percentage points over GPT-Realtime-2.1, and ranking #1 on Tau3 when paired with GPT-6 Astra at medium reasoning effort.
- In early evaluations, Speak found GPT-Live-1 cut interruptions by almost 80% versus previous turn-based systems by giving learners more time to think.
OpenAI has launched GPT-Live-1 in the API at $0.05 per minute for the front-end voice layer, making the full-duplex voice model that debuted in ChatGPT available to developers building voice-enabled apps and business workflows. The pricing covers the voice layer itself; developers pair the model with a backend reasoning model and agent harness of their choice, and pay for those separately.
The release matters because it attacks the core architectural weakness of today's voice agents. Traditional systems stitch together speech-to-text, a reasoning model, and text-to-speech, and each handoff between those stages adds latency while creating opportunities to lose timing, context, or the natural rhythm of a conversation. Developers are often the ones left coordinating those stages, including handling what happens when a user interrupts, pauses, or changes direction mid-sentence. GPT-Live-1 collapses that stack into a single model that listens and speaks simultaneously.
One model, one conversation loop
According to OpenAI, GPT-Live-1 reasons over incoming and outgoing audio together, which lets it respond to interruptions and acknowledgements as they happen rather than waiting for a clean turn boundary. Deeper reasoning and tool calls are delegated to a backend text model — OpenAI names GPT-6 Astra as one option — or to a third-party model. The conversation continues while that work happens in the background.
That split gives developers a cost and capability dial. OpenAI's documentation suggests pairing GPT-Live-1 with a model like Luna for high-volume tasks such as scheduling or order updates, and reserving a model like Astra for complex customer issues that require reasoning. The company frames this as letting developers match reasoning depth, speed, and cost to each task.
The model is not purely free-form, either. Although GPT-Live-1 is not a turn-based model, it natively supports turn detection, so developers can continue to build around explicit turn boundaries where their workflows require them. It also natively provides ASR transcripts and response text, offers what OpenAI describes as strong alphanumeric understanding, and supports keyword biasing — features aimed at the unglamorous but critical requirements of production voice systems, where correctly capturing serial numbers, names, and product codes determines whether an agent is usable.
The benchmark claims
OpenAI's own evaluations point to substantial gains over its previous real-time offering. Across the company's evaluations, GPT-Live-1 improves Full Duplex Bench performance by 30 percentage points over GPT-Realtime-2.1, with large gains in turn-taking latency and interactive behavior. Full Duplex Bench evaluates pause handling, conversational turn taking, interruptions, and backchannels.
Paired with GPT-6 Astra at medium reasoning effort, the model also ranks #1 on Tau3, a benchmark OpenAI says measures frontier voice-agent intelligence on end-to-end tasks. The accompanying evaluation suite covers several concrete domains:
- Spoken customer-service tasks in airline, retail, and telecom domains, scored by Pass@1 task success, with each domain given equal weight in the headline number.
- Spoken banking support with knowledge retrieval and account tools, where Pass@1 measures the fraction of 97 banking_knowledge tasks completed successfully.
- Reactions to background speech, speech directed to another person, listener backchannels, and interruptions.
- How quickly the agent starts its reply after the user finishes a turn.
- Tool use from spoken requests containing natural pauses, hesitations, and self-corrections, with Pass@1 scoring the tool-call sequence, using GPT Terra at low reasoning effort as the backend.
- Spoken answers to tool-using requests containing pauses and self-corrections, scored on how well the answer matches the reference intent, also with a Terra (low) backend.
These are vendor-run evaluations on vendor-defined benchmarks, so the numbers should be read as OpenAI's characterization of its own progress rather than independently verified results. Still, the size of the claimed gap on Full Duplex Bench — 30 percentage points over GPT-Realtime-2.1 — indicates how much OpenAI considers the interruption and turn-taking problem central to making voice agents feel usable.
The interruption problem, quantified by Speak
The most striking customer datapoint in the announcement comes from Speak, the language-learning company. In early evaluations, Speak found that GPT-Live-1 gave learners more time to think before the language tutor responded, cutting interruptions by almost 80% versus previous turn-based systems.
That figure gets at the practical stakes. A language tutor that barges in the moment a learner pauses is worse than useless — it removes the thinking time that learning depends on. The same dynamic applies to customer support, telephony agents, and any scenario where a user hesitates, self-corrects, or briefly speaks to someone nearby. Smooth interruption handling is listed by OpenAI as a core GPT-Live-1 strength, and the Speak result is the company's evidence that it translates into business impact.
What else the API release adds
Beyond the architecture, OpenAI has focused the API release on capabilities that let developers steer and customize voice experiences around their users, workflows, and goals. The feature list includes:
- Tone, pace, and style control. Developers shape an agent's tone, pace, and conversational style through the system prompt, rather than picking from fixed personas.
- Silent context management and background noise handling. The model handles background noise and silence without interrupting the conversation or narrating every step out loud.
- Long-session reliability. OpenAI says GPT-Live-1 improves context retention and conversational quality across extended interactions — a direct answer to the degradation that has plagued long voice sessions.
- Telephony support. The API enables deployment of full-duplex voice agents for phone calls, from restaurant reservations to customer support — a market where turn-based latency has long been the norm.
The company is also expanding voice options. OpenAI says it is moving from a small set of real-time voices to a broader selection across accents, dialects, and languages, and that it will continue to expand voice options and language availability over the coming months. Custom voice access is handled separately: developers need to contact sales to learn about eligibility and the request process.
Two paths to deployment
Developers can build on GPT-Live-1 directly through the API, or through OpenAI Presence, the enterprise agent platform the company introduced separately. Presence uses GPT-Live-1 to power real-time voice interactions and helps enterprises deploy agents that can answer questions, resolve issues, use company systems, take approved actions, and escalate to human staff when needed. Access to Presence runs through OpenAI account directors.
The direct API path is documented at OpenAI's developer site, including examples showing how an application passes conversation context to Codex and returns Codex's answer to GPT-Live-1 — a pattern for delegating agentic work while the voice conversation continues.
Why the market is watching
The voice-agent market has converged on a familiar critique: demos impress, production systems frustrate. Latency, brittle interruption handling, and context loss over long sessions have kept voice from displacing text-based chat in many deployments. OpenAI's bet with GPT-Live-1 is that the problem is architectural — that chaining STT, an LLM, and TTS cannot reproduce the timing of human conversation no matter how good each component gets — and that a single native speech model can.
At $0.05 per minute for the voice layer, plus backend model costs, the economics are pitched at volume deployments: a phone-based reservation agent or support line running full-duplex becomes a plausible build rather than a research project. The Speak result — an almost 80% reduction in interruptions versus turn-based systems — suggests the approach is already changing how users experience the technology.
OpenAI says it will keep expanding voice options and language availability over the coming months, and the pairing flexibility — Luna for volume, Astra for depth, third-party models as backends — positions GPT-Live-1 as a voice layer that other model providers' reasoning can plug into. How competitors building chained or native-speech alternatives respond, and whether independent benchmarks corroborate the Full Duplex Bench and Tau3 results, will shape whether full-duplex becomes the default architecture for voice agents.
Original: youtube.com
More from Marcus Bennett
Show full bio
Senior reporter covering consumer brands and retail at AI In Context.
108 articles
Related articles
- OpenAI ships three realtime audio models, led by GPT-Realtime-2
- OpenAI Launches gpt-realtime Speech-to-Speech Model With MCP and Phone Support
- OpenAI Kills the Turn Detector: Inside GPT-Live's Realtime Voice Architecture
- OpenAI Releases GPT-5.2, Its New Frontier Model for Professional Work
- OpenAI Launches ChatGPT Go Worldwide With GPT-5.2 Instant Access