Why Voice AI Is Still Waiting for Its ChatGPT Moment
Tech leaders from PolyAI and Otter say that despite major investment, voice AI has yet to experience a breakthrough moment like ChatGPT did for text AI, largely because unreliable speech recognition causes the entire processing chain to falter.
Major Investment, No Breakthrough Yet
Voice AI has been attracting large sums of investment for some time now, spanning model developers, enterprise customer service providers, meeting notetakers, and AI dictation software. New tools claiming to sound and converse just as naturally as a human appear on a weekly basis. Yet according to several tech executives, the sector has not yet reached the turning point that ChatGPT marked for text AI in 2022. This is according to TechCrunch, reporting on statements made last month at the HumanX conference.
Speech Recognition as the Weak Link
Shawn Wen, CTO of enterprise voice AI platform PolyAI, says the industry has indeed reached a milestone with so-called full-duplex models, which can speak and listen at the same time. But the next hurdle, he says, is reasoning speed: responses need to come faster for a conversation to feel natural. AI agents in customer service also need to avoid sounding robotic, so users can trust that their issue is actually being resolved. Wen expects that after two to three natural conversational turns with a good voice agent, customers will gradually feel less need to speak with a human.
One concrete technical problem he points to: automatic speech recognition (ASR) regularly misses important keywords in a conversation. This causes context to be lost, which makes the entire processing pipeline behind the system falter.
Otter: Transcription as a Foundation, Not a Goal
Alex Gay, CMO of meeting notetaker Otter, points to speaker identification, intent recognition, and linking these to organizational knowledge as necessary steps toward further automation. Otter is working on so-called 'digital twins' that can represent people in meetings. According to Gay, the voice of such a digital twin needs to carry the same emotional expression as a real human — otherwise the system becomes nothing more than a Q&A chatbot rather than a full conversational partner.
Gay emphasizes that transcription was never the end goal at Otter, but rather the foundational layer on which productivity gains are built. If that transcription isn't accurate enough, all subsequent actions — such as automatically generated tasks — go wrong, undermining users' trust in the platform. That's why Otter continues to invest in improving its underlying ASR model.
Trust and Transparency as Key
Both executives cite language understanding as an area where voice models still need to improve. They also agree that transparency is crucial for adoption: users need to know when they're talking to AI or when a conversation is being recorded. Wen states that companies should always make clear in enterprise conversations that the other party is an AI system. Otter aims to build trust in a similar way by informing participants that a meeting is being recorded, even when there is no physical 'bot' present in the call.
Outlook
The statements made at HumanX paint a picture of a sector making rapid technical progress, but one that, according to the executives involved, has not yet experienced a definitive breakthrough moment. Reliable speech recognition, faster reasoning, and clear transparency toward users are cited as prerequisites before voice AI can achieve an impact comparable to that of text AI following the arrival of ChatGPT.