Think about the last time you called a company's support line about a broken appliance or a device that wouldn't turn on.
Chances are you found yourself trying to decipher a painfully long serial number printed in a tiny, eye-squint-inducing font.
Or maybe you had to describe what a flashing light looked like, or explain – in words –exactly which part of the product had cracked.
I don’t think I’m going out on too much of a limb here to say that it's a frustrating experience. And according to Shan Lilja, Co-Founder of Mavenoid, it's also a completely avoidable one:
“When you're supporting physical products, voice alone is like standing next to someone but with your eyes closed. You can't see what they're doing, you can't point to anything, and they can't show you what they're talking about.”
I’m sure Shan’s words will resonate with many customers. This is the core problem multimodal support is designed to solve.
It isn’t a bells-and-whistles upgrade to existing channels; it’s more of a fundamental rethink of how resolution actually happens when the product in question exists in the physical world.
Chat Got Us Far, But Not Far Enough
The past decade of customer support innovation has been largely text-driven.
From chatbots to knowledge bases to AI-assisted ticketing, the default assumption has been that if you can get a customer to type their problem, you can get them to a solution.
For most queries, that works fine. Interactions such as account issues, billing questions, and order status updates can usually be solved through the power of the written word.
But for physical products, the gap between what a customer can describe and what a support agent needs to know is significant.
A customer troubleshooting a home ventilation system or setting up a new piece of hardware is often trying to follow fairly complicated instructions while their hands are occupied, sometimes in an unfamiliar environment, and frequently under stress.
Text creates friction at every step of that process. And friction in support compounds quickly.
The Case for Visual + Voice
Multimodal support combines voice, text, and visual elements (images, diagrams, video guidance) within a single connected experience.
The customer doesn't have to switch channels, restart the conversation, or hunt for a YouTube tutorial. Everything they need is in the same place, delivered in the format that actually fits what they're trying to do.
The practical impact of that can be considerable, as Lilja explains:
“We see 4x fewer abandoned calls with multimodal compared to text-only AI agents.”
“Take warranty claims as an example, rather than asking a customer to describe their issue over the phone, they can just take a picture.
“It's a much smoother process – much less effort for the end user – and you can automate much more of it as a result.”
Part of what drives that engagement is the removal of the small but relentless frictions that erode text-based support.
Things that might sound minor – like serial numbers being misheard over the phone, difficulty understanding product instructions, and instances where customers have to put down the device to find another device to look something up – can have a major impact on the overall customer experience.
“If someone says 'serial number SN800Z' over the phone and it's misheard, you end up repeating yourself,” Lilja explains. “Instead, it might be more appropriate to ask the user to enter the serial number via keyboard or take a photo of the label and detect it that way.

