OpenAI has released its "most advanced speech-to-speech model yet": gpt-realtime.
The AI giant also took the wraps off its Realtime API, now generally available with new capabilities as it moves out of beta.
OpenAI hopes enterprises and developers will leverage both the model and the API to build "production-ready voice agents".
Some of the API’s latest features will help here. For instance, the new ability of the Realtime API to support image inputs and remote MCP servers will make these agents more capable.
Yet, there’s also an exciting capability to develop better customer support voice AI agents.
As Peter Bakkum, Member of Technical Staff at OpenAI, said in the announcement video:
We’ve added support for SIP telephony, which makes it much easier to build applications for voice-over-phone situations like customer support.
With this, a developer could easily grab a phone number from Twilio, feed that into the SIP interface provided by OpenAI, add prompts, feed it data, and let it go.
As Andreas Granig, CEO at Sipfront, observed in a LinkedIn post, that is quite the threat to many conversational AI startups.
"There are quite some startups, who only provide an interface to the public phone network for existing speech-to-speech AI services, often without much telco moat, but relying mostly on Twilio and others… They are in hot water now," noted Granig.
The CEO acknowledged that startups specializing in tool calling for advanced integrations remain safe, since that remains a specialist field. However, he added:
The voice interface for AI assistants just became [a] commodity.
As a result, it will be more difficult to differentiate use cases for AI assistants, signalling to many conversational AI startups that now is the time to step up.
What About the New gpt-realtime Model?
OpenAI hopes many customer support teams will leverage gpt-realtime, alongside the Realtime API, as they advance their customer support automation strategies.
Indeed, as Peter Bakkum, Member of Technical Staff at OpenAI, said in the announcement video:
We carefully aligned the model… to real scenarios like customer support and academic tutoring.
There are many reasons why support leaders would consider the gpt-realtime model. For starters, it enables AI agents that can understand and produce audio without relying on separate transcription, language, and voice models.
Additionally, there are performance benefits. For instance, these agents will respond faster, as it’s just one model, and capture subtleties like laughter or sighs while expressing various emotions.
OpenAI also claims the model can deliver more natural, high-quality audio while following instructions across complex, multi-turn conversations.
Developers can also adjust pace, tone, style, and even roleplay characters.




