Last week, OpenAI released the latest iteration of its flagship large language model (LLM): GPT-4o.
The LLM stands apart in its multimodality, with the ability to reason in real-time across audio, vision, and text.
For years, building AI that understands multiple modalities has proved challenging. Just creating pipelines for tasks such as speech-to-text is tricky due to issues like high processing time.
Now, GPT-4o can do that almost instantly.
Yet, until this point, AI platform providers have invested heavily in such multimodal projects, funding that Rowan Trollope, CEO of Redis, now suggests is obsolete.
Taking to X (formerly Twitter), the once Five9 CEO stated:
Hundreds of millions of dollars of R&D in Contact Center AI Agents just became obsolete with OpenAI GPT-4o developments. The attention should be on back-end automation
It would not be the first time ChatGPT has made R&D into contact center innovation obsolete, either.
Just consider how earlier releases of the LLM have undone the work of many analytics providers, which spent hundreds of hours engineering natural language processing (NLP) models to gauge intent, sentiment, and more. A LLM can extract all this out of the box.
Another example is agent-assist innovation. While businesses once spent significant R&D resources building use cases like isolating key data points within a customer conversation, ChatGPT and other LLMs can do so instantaneously.
Indeed, that is ultimately why many businesses look to first implement LLMs in the contact center. The use cases were already there, but they are now far more accessible.
Consider the following graphic from an October 2023 Gartner study. It showcases how customer service is a prominent recipient of enterprise GenAI investment.

Expect this to continue, and - with the introduction of multimodal LLMs - the use cases will become more innovative. Real-time translation is an excellent example.
Real-Time Translation and Other Multimodal Use Cases
In recent years, conversational AI vendors have brought various real-time translation models to market, with brands like Cognigy even making them available on the voice channel.
Typically, these apps will first use speech-to-text to develop a transcript from the customer’s audio.
That transcript then feeds through a translation engine – like Google Translate – and the agent receives a text translation within their workspace.
From there, the agent types out their reply, which - via the engine – translates back to the original language and plays out through a text-to-speech audio stream.
The primary issue with these experiences is the dead air between the customer speaking and the agent typing back a response. It’s a rapport killer.
Thankfully, with “off-the-shelf” real-time translation, a multimodal LLM may cut to the chase.
Just take a look at this example that OpenAI released of GPT-4o, translating a live conversation between native English and Spanish speakers.
Alongside translation, consider how AI can adjust the agent’s accent to one that’s much more familiar to the customer, ensuring full comprehension.
Krisp already offers this use case. Yet, with GPT-4.0, it may become much more widely available. After all, one of OpenAI's demos showed a GPT-powered voice changing styles on the fly.
