Text-based chatbots got their breakthrough moment with ChatGPT in late 2022. Voice, by contrast, is everywhere and nowhere: it sits inside phones, cars and call centres, yet no single product has made talking to software feel inevitable. That is the assessment several industry executives gave to TechCrunch, arguing that voice AI has still not found its equivalent moment.
The gap is not about model quality. Modern speech recognition, text-to-speech and conversational models are individually good enough for many tasks. The problems are structural. A voice interaction is a chain — speech-to-text, a language model, then text-to-speech — and every link adds delay. Once the total round trip stretches past roughly a second, conversation stops feeling like conversation and starts feeling like a walkie-talkie.
Two harder problems sit on top of latency. The first is turn-taking: knowing when a user has genuinely finished speaking, and handling interruptions without talking over them. The second is recovery. A voice agent that fails five per cent of the time is mildly annoying in a chat window; over the phone it can strand a caller with no obvious way to go back.
Where voice does work today, it tends to be in narrow, high-volume lanes: call-centre triage, order taking, clinical note-taking and in-car assistants. Those deployments succeed because the vocabulary is bounded, the failure path usually ends with a human, and the return on investment — minutes saved per call — is easy to measure.
Executives argue the breakthrough will come from reliability and cost, not from another polished demo. Sub-second latency, dependable interruption handling, sensible escalation to a human and per-minute pricing that works at scale are what turn a laboratory trick into infrastructure. Until then, voice remains a feature bolted onto products rather than a platform in its own right.




