Imagine a customer asks your chatbot if their refund will arrive by Friday. The bot calculates the odds and replies that it is ‘likely’.
To the bot, ‘likely’ might mean an 80% statistical chance. But to the anxious customer, ‘likely’ often sounds like a soft ‘yes’. If the refund does not arrive, the customer does not feel the bot made a statistical error. They feel misled.
This scenario highlights a subtle risk in chatbot language. It is not about hallucinations or wrong answers. It is about estimative uncertainty.
New research suggests that Large Language Models (LLMs) and humans interpret words of probability very differently. For CX leaders, this ‘translation gap’ represents a potential friction point that could be eroding trust.
The Mathematics Of Misunderstanding
The core of the issue lies in how we assign numbers to words. A recent study titled ‘An evaluation of estimative uncertainty in large language models’ explored this dynamic. The researchers compared how LLMs interpret probability words against human benchmarks.
The results highlighted a distinct mismatch.
When a human hears the word ‘likely’, they might internally calibrate that to a 65% chance. But the study suggests that LLMs can assign a significantly higher probability to the same term, often pushing above 80%.
This gap might seem small mathematically. In a customer service context, however, it is potentially massive. It could be the difference between managing expectations and setting a customer up for disappointment.
If your automated agent uses confident language to describe uncertain outcomes, it risks overpromising. The bot isn't lying. It is simply speaking a different statistical dialect than your customer.
Mayank Kejriwal, Research Associate Professor at University of Southern Carolina summarizes the research in an article on Fortune:
"An AI model might use the word ‘likely’ to represent an 80% probability, whereas a human reader typically interprets it as closer to 65%."
Context Changes Everything
The risk becomes more complex when you factor in context. The study indicates that LLMs are highly sensitive to how a prompt is phrased.
Changing the language of the prompt or the framing of the question can shift the bot's probability estimation. A bot might interpret ‘likely’ differently in a financial context versus a casual conversation.
This variability makes it difficult for conversation designers to guarantee a consistent experience. A human agent knows that telling a banking client "funds will ‘likely’ clear" carries more weight than telling a shopper "this shirt will ‘likely’ fit."

