Key facts
- Voice AI has not yet reached its 'ChatGPT moment,' according to PolyAI CTO Shawn Wen.
- Full-duplex voice AI models can speak while listening, but fast reasoning is needed for natural conversations.
- Otter is developing digital twins for meetings that require output voices to have the same emotive expressions as humans.
- Improved transcription accuracy is critical for voice AI tools to maintain user trust.
- Transparency is key, with tools needing to disclose AI interaction or recording.
Despite significant investment and rapid development in voice AI, the technology has not yet achieved a breakthrough moment comparable to ChatGPT, according to industry executives. Shawn Wen, CTO of PolyAI, noted that while full-duplex models, which allow for simultaneous speaking and listening, have been developed, the next critical step is to significantly speed up reasoning capabilities. This will enable AI models to fetch answers quickly, making conversations feel more natural and less robotic.
Wen believes that for AI agents in customer service, sounding natural and instilling confidence in callers is paramount. He suggested that as voice AI improves to a point where customers are willing to engage for several turns, they may become comfortable relying on AI for problem resolution, potentially reducing the need to speak with a human agent.
Alex Gay, CMO for Otter, a meeting notetaker company, highlighted the importance of accurate speaker identification, intent capture, and integrating organizational knowledge to enable automation. Otter is also exploring digital twins for meetings, emphasizing that the AI's voice must convey the same emotive expressions as a human to facilitate meaningful debate and strategic discussions, rather than just acting as a question-and-answer chatbot.
