From an organizational perspective, the case for multimodal AI agents in CX is fairly open and shut.
Multimodal can enable richer, visual, voice-enabled interactions that lead to better outcomes, reduce customer effort, and build the kind of trust that text-only AI consistently struggles to earn.
So why are some organizations still seemingly wary of adopting the technology?
For Shan Lilja, Co-Founder of Mavenoid, it's more about familiarity than capability:
“Companies worry about how customers will react to multimodal AI agents. But for customers, it's a far more natural experience.”
While he concedes that this can be a little daunting for companies, he explains that the full multimodal leap doesn't have to happen all at once.
Multimodal adoption is a journey with distinct phases, each one delivering real value independently, and the first step is considerably more accessible than most organizations assume.
Start Where the Difference Is Obvious
The instinct when approaching a new capability is often to either go all in or hold back entirely. Lilja's advice is to do neither.
Instead, he encourages organizations to “design a set of steps so that each step is a no-brainer.”
That means beginning with support scenarios where multimodal is not marginally better than text, but, according to Liljas, can deliver “10x better” moments.
Strong use cases for multimodal include:
- Troubleshooting a complex physical product
- Onboarding a customer through a multi-step hardware installation with their hands occupied
- Processing a warranty claim where sharing images or video requires far less effort than a verbal description
These are the use cases where the case for multimodal is so self-evident that adoption becomes much easier to drive for both the brand and the customer.
The Four Phases
From there, Lilja maps out a progression that takes organizations from their current state to the frontier of what's now possible.
Phase one is channel switching, which he refers to as “pre-multimodal.” Voice automation detects that a visual element would help and sends the customer a text link with an image or guide.
The two modalities aren't unified, but they're connected. It's a meaningful first step that improves the customer experience.
Phase two brings those modalities together. Voice, text, and visual inputs all flow into a single session, allowing the customer to talk, type, and share images without switching context.
“Friction is much lower,” says Lilja. This is where the Broan-NuTone results were largely delivered: call drop rate fell from 25% to 7%, speed of answer improved fourfold, and service level score almost doubled from 43% to 80%.
As Don Lyskawa, Product Support Manager at Broan-NuTone, put it:
“The combination of digital self-service and voice automation has significantly reduced pressure on our team and improved the overall experience of our customers.”
Phase three is what Lilja calls being “hardcore multimodal” – a unified communication experience where customers can interact through any combination of pictures, voice, and text, with the AI orchestrating across all of them in real time.

