OpenAI Launches gpt-realtime Speech-to-Speech Model With MCP and Phone Support
OpenAI ships gpt-realtime, a more advanced speech-to-speech model, plus Realtime API updates: MCP server support, image input, and SIP phone calling for voice agents.

Updated
Why it matters
- OpenAI released gpt-realtime, a more advanced speech-to-speech model, alongside Realtime API updates.
- The Realtime API now supports MCP servers, image input, and SIP phone calling.
- SIP support lets developers deploy voice agents on traditional telephony networks, not just web-based interfaces.
OpenAI has released gpt-realtime, a more advanced speech-to-speech model, alongside a set of Realtime API updates that include MCP server support, image input, and SIP phone calling.
The announcement, titled "Introducing gpt-realtime and Realtime API updates," positions the new model as a step up from OpenAI's previous speech-to-speech offering. Speech-to-speech architectures skip the traditional pipeline of transcribing audio to text, processing the text, and synthesizing a response. Instead, they map audio directly to audio, which reduces latency and preserves vocal nuance that text-based transcription strips out.
The move matters for developers building voice agents and real-time assistants. Voice has become one of the most competitive fronts in AI deployment, and the quality of the underlying speech model often determines whether an agent can handle interruptions, accents, and natural conversational rhythm without awkward pauses or transcription errors.
The three accompanying API capabilities extend where and how developers can deploy the model.
MCP server support brings the Model Context Protocol into the Realtime API. MCP is an open standard that lets models connect to external tools and data sources in a standardized way. With server support in the Realtime API, voice agents built on gpt-realtime can call out to MCP-compatible services mid-conversation, potentially retrieving information or triggering actions without leaving the audio interface.
Image input adds a visual channel to what has been an audio-only API. A developer building a support agent, for example, could let a user share a photo during a live voice session and have the model reason over both the image and the spoken conversation at once.
SIP phone calling support connects the Realtime API to the traditional telephony infrastructure. SIP, the Session Initiation Protocol, underpins most business and carrier phone systems. Native SIP support means developers can build agents that answer and place calls on ordinary phone networks, not just over web sockets in an app — a prerequisite for any serious deployment in customer service, outbound calling, or phone-based assistance.
Taken together, the updates push the Realtime API beyond web-based voice chat toward the full stack of channels where spoken interaction actually happens: browsers, apps, and the phone system, with tool access and vision layered on top.
The competitive context is direct. Google, Amazon, and a crop of voice-focused startups are racing to power real-time conversational agents, and telephony integration in particular has become a battleground for enterprise AI calling products. OpenAI's decision to support SIP natively signals it wants its models answering real phone lines, not just demo interfaces.
For enterprises, MCP support reduces integration friction: rather than writing bespoke tool connectors for each deployment, developers can point the model at existing MCP servers. That standardization pressure is one reason MCP adoption has accelerated across the industry since its introduction.
OpenAI has not stopped at model quality alone. The pairing of a stronger speech-to-speech model with infrastructure-level features — protocol support, multimodal input, telephony — suggests the company is competing on the completeness of the real-time stack as much as on benchmark performance.
The rollout gives developers an immediate question to answer: whether the new model's speech-to-speech quality and the added channels justify migrating existing voice deployments, or building new ones, on OpenAI's stack rather than assembling components from multiple providers.
Source: OpenAI News
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
122 articles
Related articles
- OpenAI ships three realtime audio models, led by GPT-Realtime-2
- OpenAI puts GPT-Live-1 in the API at $0.05 per minute
- OpenAI Kills the Turn Detector: Inside GPT-Live's Realtime Voice Architecture
- Google Ships Upgraded Gemini 2.5 Flash Native Audio and Live Translation
- Google's Flash TTS models build AI voices from text descriptions