OpenAI Kills the Turn Detector: Inside GPT-Live's Realtime Voice Architecture
GPT-Live drops the turn detector for a full-duplex voice model, with Go media paths, WARP's six-to-one round-trip cut, and silent production testing behind ChatGPT Voice.

Updated
Why it matters
- GPT-Live is a full-duplex third-generation voice system that removes the turn detector from the audio path and can delegate to frontier models like GPT-5.5 asynchronously
- Rewriting the media frontend and inference logic in Go brought the new system's p95 frame-delivery latency down to the previous Python asyncio system's p50
- WARP, an open IETF proposal backed by libwebrtc and Pion, cuts WebRTC media and data startup from six network round trips to one; with Instant Connect, a session starts with a single UDP packet
OpenAI has removed the turn detector from the audio path entirely. Its third-generation voice system, GPT-Live, uses a full-duplex voice model that can listen and speak at the same time, eliminating the component that previous voice AI systems relied on to guess when a user had finished speaking. When deeper reasoning or tool use is needed, GPT-Live consults frontier models such as GPT-5.5 without interrupting the flow of the conversation.
The engineering stakes are straightforward. In older architectures, small turn-detector models faced what OpenAI describes as an unenviable task: "guess too soon, and the user gets cut off; guess too late, and the response feels sluggish." Only after the detector decided could the larger LLM start working. Human speakers hand off to each other in a fraction of a second, and turn-based systems could not match that rhythm. OpenAI says it reworked model inference, context management, and media transport over the last six months to close the gap.
The payoff goes beyond snappier conversations. The same architecture now powers capabilities in ChatGPT Voice including the newly launched ability to control a computer and coordinate agents in the ChatGPT desktop app, and OpenAI says it will underpin an upcoming GPT-Live API.
From turn taking to streaming
Earlier voice architectures inherited the turn-based nature of text LLMs, with each turn represented as a discrete audio blob. Cascaded systems ran speech-to-text, the LLM, and text-to-speech in series. That sequencing added latency and discarded cues such as tone and pacing. Speech-to-speech models improved responsiveness by processing audio directly, but the turn detector still gated when inference could begin.
GPT-Live inverts the design. Audio flows continuously in and out of the voice model while deeper reasoning and tool use happen asynchronously. The system's primary job, in OpenAI's words, is to "sustain an uninterrupted media loop." Everything else — invoking frontier models, persisting the conversation — happens off the live path.
That separation has a practical consequence: applications can change tools, policies, and backend behavior without touching the media frontend responsible for real-time audio. The live path stays small and predictable.
Go, WebRTC, and the media fast path
OpenAI made an early decision to split media flow from application and business logic. Audio moves between client and voice model on a dedicated fast path; delegation, tool use, and other application work sit behind an asynchronous RPC boundary. A slow tool call can delay its own result but cannot stall the flow of media.
The team wrote the media frontend and inference logic in Go, replacing a previous Python asyncio implementation. The rewrite significantly improved frame delivery smoothness: the new system's p95 latency matches the previous system's p50.
WebRTC provides the transport foundation. It survives packet loss, clock drift, and connection changes, and can subtly stretch late-arriving audio to prevent gaps, then briefly accelerate playback to return to real time. Minimizing buffering and blocking throughout the stack delivers what OpenAI calls the sub-second responsiveness humans expect from conversation.
Stateful inference without interruptions
Long-running voice sessions create two problems: context grows continuously, and model instances spin up and down with demand. OpenAI's answer is a seamless handoff mechanism. The system warms a replacement model instance alongside the existing one, prefills it with the current session context, runs inference on both in parallel, and cuts over when the new instance is ready.
The same mechanism handles context compaction. Compaction changes past context and therefore invalidates the model's key-value cache, which would normally force a costly re-prefill. Instead, while the original instance keeps chatting, the system compacts the context and prepares a replacement instance, then switches over without any media interruption. This lets the system support long-running calls, compacting whenever necessary.
Talking and thinking as two models
GPT-Live's delegation to frontier models decouples talking from deeper thinking, but the voice model can only mask a slow response for so long. OpenAI therefore treated the full delegation loop — routing, prompt processing, inference, and tool calls — as part of the responsiveness budget.
The first optimization happens before delegation is even requested: when a voice session starts, the application server creates an inference session for the frontier model and prefills it with the initial conversation context. That session stays warm for the duration of the conversation, backed by stable session affinity and prompt caching. OpenAI also tuned reasoning effort, output limits, tool schemas, and model-tool round trips to speed up useful results.
A second problem: the voice model operates on continuous speech, but ChatGPT's conversation UI, analytics, and safety infrastructure still operate on discrete user and assistant turns. The application server therefore teases overlapping, occasionally ambiguous speech into a queue of messages using partial transcripts and timing signals. The newest message stays provisional — its text, timing, and speaker assignment can change — until a speaker has held the floor long enough for reliable attribution.
Speaker overlap complicates the logic. A brief "mm hmm" from the assistant while the user talks should not become its own message; a substantive interjection often should. OpenAI frames the tradeoff plainly: "Every segmentation policy trades freshness for certainty." The system maintains two views of the conversation — a speculative view for the UI and an authoritative record for the analytics pipeline.
One UDP packet to start a session
Responsiveness begins the moment the user clicks the button, which puts the entire startup sequence on the critical path. Vanilla WebRTC requires a surprising number of handshakes and round trips, partly because its underlying protocols repeat work — OpenAI notes each protocol included its own anti-DoS mechanism even when the full stack didn't need it.
The company's answer is the WebRTC Abridged Roundtrip Protocol (WARP), which cuts media and data startup from six network round trips to one. WARP combines backward-compatible improvements: piggybacking the DTLS handshake over ICE (SPED), the faster DTLS 1.3 handshake, pre-negotiating the SCTP handshake (SNAP), and pre-negotiating data channels instead of using DCEP. OpenAI developed WARP as open specifications with the WebRTC community and is advancing the proposals through the IETF's TSVWG working group. WARP support already exists in libwebrtc and Pion, with efforts underway in other implementations.
A second optimization, Instant Connect, removes the SDP signaling exchange from the critical path by negotiating parameters ahead of time without reserving server capacity or requiring changes to existing WebRTC implementations. If the pre-negotiated parameters are valid, the server materializes the session when the first media packet arrives; if stale, the standard signaling flow is already underway as a fallback. Combined, the client can start a session with a single UDP packet.
Lessons from silent production testing
A system can look fast on paper and still stall under real traffic. Before launch, OpenAI routed a gradually increasing share of production ChatGPT Voice sessions to both the existing Advanced Voice Mode and the new system, with the shadow path running inference in read-only mode. Users heard no difference.
The test reshaped the company's capacity model. Voice sessions stay open and stream frames continuously, so CPU-side handlers, queues, and network paths must scale alongside inference. A supporting component saturated earlier than load tests predicted, causing requests to accumulate and latency to compound. OpenAI reframed the question from "How many requests can a GPU handle?" to "How many concurrent sessions can the system sustain while keeping every frame on schedule?"
Geography became a first-order concern; routing sessions to distant capacity added delay at several points, and the team began validating rollouts alongside regional capacity and traffic-steering configuration. Long-running sessions exposed memory pressure, reconnects exercised compaction and state restoration, and ordinary disconnects revealed races in the shutdown handshake — failures that short load tests missed because they depended on time and accumulated state. The team responded with granular telemetry, staged ramps, and the ability to isolate or disable individual paths quickly.
OpenAI now describes the GPT-Live architecture as a broader platform for realtime interaction, one that will carry ChatGPT Voice from conversation into agentic coordination and, via the upcoming GPT-Live API, let voice experiences span more devices, apps, and modalities without losing the immediacy that makes them feel live.
Original: cdn.openai.com
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles
Related articles
- OpenAI puts GPT-Live-1 in the API at $0.05 per minute
- OpenAI ships three realtime audio models, led by GPT-Realtime-2
- OpenAI Launches gpt-realtime Speech-to-Speech Model With MCP and Phone Support
- OpenAI Launches GPT-5, Claims State-of-the-Art Results Across the Board
- OpenAI Releases GPT-5.2, Its New Frontier Model for Professional Work