PolyAI CTO: Voice AI Still Waiting for Its ChatGPT Moment
PolyAI CTO Shawn Wen says voice AI hasn't reached its "ChatGPT moment" despite full-duplex models; fast reasoning, accurate ASR and caller trust remain the blockers.
Updated
Why it matters
- PolyAI CTO Shawn Wen says voice AI has not reached its ChatGPT moment despite full-duplex models being released.
- Investors have poured billions of dollars into voice AI startups, from model makers to enterprise customer service providers.
- Otter CMO Alex Gay said transcription was never the end point, and flawed ASR undermines all downstream automation and user trust.
- Otter is developing digital twins that could represent people in meetings, requiring emotive voice output.
- Both executives called for transparency: declaring when calls are recorded or when users are talking to AI.
PolyAI CTO Shawn Wen says voice AI has not yet reached its "ChatGPT moment," even after the industry released full-duplex models that can speak while listening to a user. His verdict, delivered on stage at the HumanX conference last month, cuts against the current wave of product launches: every week brings a new model or tool claiming to sound human and converse like one, and Wen argues that reality does not match the claims.
The stakes are considerable. Investors have poured billions of dollars into voice AI startups across the stack — from model makers to enterprise customer service providers, from meeting note-takers to AI-powered dictation tools. If the underlying technology still falls short of natural conversation, that capital is betting on an interface shift that has not fully arrived.
What is still missing from voice AI?
For Wen, the gap is not in the voice itself but in the speed and quality of reasoning behind it.
"We have reached the milestone of developing full-duplex models. The next challenge is to make reasoning very fast, so that the models can fetch answers quickly and the conversation feels natural," he said at HumanX.
Full-duplex capability — the ability to speak while listening, as humans do — was long treated as a milestone for conversational AI. Wen's point is that it is a necessary condition, not a sufficient one. A voice that overlaps speech naturally still fails the caller if the model cannot retrieve and compose an answer quickly enough to keep the exchange feeling fluid.
He added a second requirement for enterprise deployments: AI agents in customer service should not sound robotic, and they must give callers enough confidence that the agent can actually solve their problem.
Wen sketched a progression for how that confidence builds. "I think the next stage will be slightly different because once the voice is good enough, like, and the customer is willing to engage with them for the first two or three turns, they start to build confidence, and over time, they will feel like I probably don't have to talk to a human if the agent can solve my problem," he said.
That framing matters for enterprises measuring containment rates — the share of calls resolved without a human agent. On Wen's reading, caller trust is earned turn by turn: the first two or three exchanges determine whether a customer stays on the line with the AI at all.
What does accurate transcription unlock?
Alex Gay, CMO of meeting notetaker Otter, pointed to a different bottleneck: speaker identification, intent capture, and combining that with organizational knowledge. He described these as key steps for enabling automation beyond the transcript itself.
Otter is also developing digital twins that could represent people in meetings. For that technology, Gay said the output voice must carry the same emotive expressions a person would hear when talking to a human in a meeting.
His standard for what counts as success is demanding. "If you think about the meetings that you're in right now, the best conversations that you have are where you can have debate, and strategic discussions, and when you feel like there's a relationship that underpins it. If you aren't able to have that with an avatar, then it's just a q and a chatbot," Gay said.
The distinction matters commercially. A chatbot that answers questions is a commodity; an avatar that can sustain debate and relationship-driven discussion would open far more use cases for automated meeting attendance.
How big is the speech recognition problem?
Both executives agreed that the foundation layer — automatic speech recognition, or ASR — remains a weak link despite improvements in voice AI models. AI assistants often misunderstand users, and meeting notetakers still produce wrong transcripts or wrong summaries.
Wen said ASR models often miss important keywords, which creates a problem in capturing the full context of a conversation. A misheard word does not just corrupt one line of text; it distorts everything downstream built on it.
Gay connected the issue directly to trust and to Otter's product strategy. "For Otter, you know, transcription was never the end point. It was just the layer that we could start to drive some of the productivity gains on the back of. But if your original transcription didn't have the accuracy that you needed, all of the follow-up actions that you have become flawed. And the minute that starts to take action, that is wrong. You lose trust in the platform. It is critical for us to continue to improve that ASR model because all of the downstream impacts are significant," he said.
He also flagged language coverage as an area where voice models still need to improve, and said Otter keeps working on improving transcription.
The pattern Gay describes is the core risk of the current voice AI stack: each layer of automation compounds the errors of the layer beneath it. Transcription errors propagate into summaries, summaries into follow-up actions, and follow-up actions into user trust — which, once lost, is hard to rebuild.
Should AI disclose itself?
The new generation of voice tools also raises transparency questions that predate the technology's maturity. Gay argued that tools should declare to customers that they are being recorded or that they are talking to an AI.
Otter's approach extends beyond its own bots. The company wants to instill trust in everyone in a meeting, so even for meetings where the Otter bot is not present, it wants to try methods such as notifying everyone in the chat that the meeting is being recorded. Wen agreed on the enterprise side, saying it is important to establish that people are talking to an AI on enterprise calls.
Disclosure is not an abstract policy debate here. PolyAI builds voice agents that answer customer service calls, and Otter's bots attend meetings where participants may not know an AI is present. As these agents improve — faster reasoning, better voices, digital twins — the moment a caller or meeting participant realizes they are talking to software becomes the point at which trust is either established or lost.
What comes next?
The two executives' comments sketch a concrete roadmap for the industry: faster reasoning to make full-duplex conversation feel natural, more accurate ASR to protect downstream automation, emotive voice output strong enough to sustain real debate, and disclosure practices that keep users informed.
Until those pieces land, the billions invested in voice AI are funding a bet on an interface shift whose "ChatGPT moment" — in Wen's phrase — still lies ahead.
Source: TechCrunch AI
More from Elena Vasquez
Show full bio
Market editor covering media and advertising at AI In Context.
223 articles
Related articles
- ElevenLabs launches Eleven v4 with tighter voice control
- OpenAI ships three realtime audio models, led by GPT-Realtime-2
- Google Ships Gemini 3.8 Live and Extended Thinking Models
- Alibaba's Qwen launches Audio 3.1 with five models, cuts prices up to 95%
- OpenAI puts GPT-Live-1 in the API at $0.05 per minute