I can remember experimenting with early offerings in speech-to-text software years ago, expensive applications which required hours of practice and training to use — but then, self-defeatingly, also required hours of editing the text afterwards, to turn the copy into anything meaningful or usable. As Deepgram CEO, Scott Stephenson, reflected, it wasn’t until really recently that the accuracy level finally crossed the threshold for the technology to be of real utility, breaking out first in the domestic environment.
“Business decision makers were seeing their kids using Alexa to do their math homework for them, and suddenly saying wait a minute… We need to take another look at this.” So while the tech powering general speech recognition in home devices is not exactly the same as that used in business devices, it powerfully illustrates the potential for new ways of driving user interactions.
Driven by voice: New ways to get things done
[caption id="attachment_30120" align="alignright" width="200"]
Scott Stephenson[/caption]
Deepgram’s end-to-end deep learning method achieves unprecedented accuracy levels using ‘command and control’ type inputs, which are very different to the conversational type of speech recognition general devices work with. “It's capturing audio at a much higher sampling rate, and storing it in a lossless format and then sending it over the Internet to be transcribed by dedicated servers, using a speech recognition system that is anticipating command and control to come to it. It’s expecting you to say things about maps and taking you places and addresses and things like that, or give commands like turn on the lights — it's not expecting you to have a two-way conversation with another human for an hour, that would be out of domain for it.”
It’s this ability to learn from very specific inputs, as well as adapting to individual acoustic environments, which enables their model to exceed 80% accuracy rates within a couple of weeks.




