The pilot worked. The production deployment did not. For most enterprises, the difference comes down to one thing they did not plan for: the voice layer.
It has never been easier to build a voice AI agent. LLMs, text-to-speech, and speech-to-text have matured to the point where a convincing proof of concept is achievable in days. The problem is what comes next. Connecting that agent to real telephony infrastructure, backend systems, and human escalation workflows has not been accelerated by generative AI. That work is still hard, and it is where the gap opens.
Yehuda Herscovici, VP of Product at AudioCodes, has watched this play out repeatedly:
"It became almost trivial to design and have a very impressive voice AI agent, to start pilots, to do POCs. But then, once you are happy with the voice AI agent you developed, you need to put it into production. This is where you need to integrate it with your backend systems, with the knowledge base, with the telephony systems."
The compliance work, the telephony stack integration, the escalation design: none of that moves faster because the LLM got better. Deployments stall not because the AI failed, but because everything around it was not ready.
Telephony Integration Is Harder Than It Looks
In a pilot, you provision a phone number and test in isolation. In production, you integrate with what the enterprise already has: SIP trunks, on-premise contact centres, UC platforms, all built on VoIP protocols with inconsistent implementations. The voice AI stack sits in an entirely different architectural world.
"To bridge between these two worlds, also to do it at scale and with good voice quality and redundancy, is very hard," says Ilan Avner, Director of Product Management at AudioCodes.
The problem is compounded for organisations mid-migration, running on-premise infrastructure while moving toward cloud platforms. Most do not realise their voice AI investment may not survive that transition until they are already committed to it.
The Voice Layer Is Invisible Until It Fails
There is an infrastructure problem that rarely gets discussed until it causes a deployment to fall over. If the audio quality at the transport layer is poor, the AI stack built on top of it will fail. When that happens, the blame lands on the AI.
"If you do not feed a voice AI agent with high-quality voice at the transport and infrastructure layer, it will fail. If it works well, nobody talks about it. But if it fails, everybody blames the AI and says it is not ready for prime time."
Herscovici's analogy is electricity: a data centre can be world-class, but without a reliable power supply, nothing runs. Noise filtering, latency control, and voice clarity at the transport layer are not refinements. They are the foundation. AudioCodes brings 30 years of VoIP infrastructure experience to this layer, which is central to how both Live Hub and Voice AI Connect are positioned.

