Microsoft has added a new feature to its Dynamics 365 Contact Center platform, Constrained Speech Recognition.
The innovation introduces structured rules to increase the accuracy of voice inputs.
As more contact centers apply AI tooling to the voice channel, such as conversational analytics, agent assistance, and automation, speech recognition engines are playing an increasingly crucial role in supporting the success of these implementations.
However, traditional voice recognition systems can struggle to accurately understand what customers say, because they are designed to interpret a wide range of possible words without focusing on the specific context and intent of the conversation.
Human agents naturally use contextual cues, including the subject of the call, related common phrases, and tone of voice to anticipate and understand what the customer is likely to say. They can also account for accents, slang, muffled speech or unexpected wording more easily than an automated system.
Constrained Speech Recognition aims to close the gap. It uses structured rules known as "grammars" to define what the system should recognize, and help narrow down the words and phrases the customer is likely to use to reduce errors.
Grammars typically use the Speech Recognition Grammar Specification ("SRGS") format, which is an industry standard that can include logic for validation, positional constraints, and checksum verification. This is key in sectors like healthcare, finance, and enterprise IT, where a misheard word or number can disrupt the customer experience.
Additionally, grammars can help voice recognition systems recognize when a user is citing an alphanumeric string like an ID number, confirmation code, or package tracking reference. It can also help identify items from a specific list.




