Speech-to-text technology (also known as automatic speech recognition) is not a new concept in the contact center landscape. In the 1950s and 60s, innovators were already using speech recognition systems to identify people speaking digits out loud. These solutions helped to inspire the creation of countless future CX tools.
Now, speech-to-text models are used in everything from IVR workflows, to payment processing, conversational analytics, and self-service systems. As voice data continues to be one of the most important resources in a contact center, speech-to-text technology have grown increasingly more advanced, leveraging new forms of natural language processing, understanding, and machine learning.
The question is, how accurate are speech-to-text solutions in 2023?
The Importance of Accuracy in Speech-to-Text
Phone conversations are still the main way businesses and consumers interact. However, manual conversation analysis requires significant time and effort. Speech analysis software, leveraging automatic speech recognition (ASR) technology and speech-to-text can alleviate these problems.
Speech-to-text technology can be used for everything from the creation of powerful AI chatbots, to discovering business insights and intelligence. However, the value of any speech-to-text offering hinges entirely on its accuracy. Over the years, voice assistants and AI tools have been able to achieve higher levels of accuracy, thanks to improved algorithms and data sources.
“The Speech AI and generative AI tools built on top of audio and video data today are only as useful as the accuracy of the conversational data being fed into them. If the transcript isn’t accurate, the insights and suggestions the AI models output won’t be accurate either. Highly accurate speech-to-text is critical for building a strong foundation,” says Prachie Banthia, VP of Product at AssemblyAI.
However, despite the sophisticated nature of AI solutions for processing human language today, many speech-to-text systems still suffer from various accuracy issues. Benchmarks published in 2021 found that Amazon’s speech-to-text technology still had an error rate of 18.42%, Microsoft’s error rate ranged at 16.51% and Google video was ranked at 15.82%.
Ultimately, no speech-to-text solution is 100% accurate. All of these systems encounter limitations, whether it’s struggling to understand accents or dialogues, or being unable to distinguish quiet speech from large amounts of background noise.
Notably, speech-to-text AI models can also only recognize words in their existing database. This means it’s possible for some models to overlook certain words and phrases entirely.
Measuring Speech-to-Text Accuracy: WER
Since the value of any speech-to-text solution depends on its accuracy, companies creating intelligent systems need to regularly test and evaluate their systems for potential discrepancies. The most common way of measuring the accuracy of a speech-to-text solution, is Word Error Rate, or WER.
Word Error Rate calculates the exact number of errors present in a transcription produced by an ASR system, compared to a human transcription. It’s the de-facto standard for measuring how effective an automatic speech recognition tool, or text-to-speech solution is.
However, WER may not be the only metric worth examining when analyzing the utility of a speech recognition system. Word Error Rate can tell companies and developers how different an automatic transcription is to a human transcription. However, WER isn’t smart. It can’t account for context and legibility, but instead only looks at substitutions, deletions, and insertions in texts.




