Uncertainty estimation in LLMs: the reasoning language matters

Discover how reasoning in English improves uncertainty in multilingual LLMs. Practical tips for choosing the method according to scale.

miércoles, 8 de julio de 2026 • 2 min read • Q2BSTUDIO Team

Reasoning in English improves uncertainty in low-resource languages

Uncertainty estimation in large language models (LLMs) has become a key piece for deploying reliable artificial intelligence systems, especially when the model needs to recognize when it should refrain from answering. Until now, most research focused on English, leaving aside the enormous challenge of operating in a multilingual world. A recent study evaluated nine uncertainty estimation methods in 22 languages, revealing a finding that changes the perspective: the language in which the model internally reasons has more impact on reliability than the language of the original question. When the LLM is forced to reason in English, even for questions in low-resource languages, uncertainty is significantly reduced and the performance gap between languages is closed. This suggests that understanding minority languages is not the real problem, but rather that the bottleneck lies in the ability to generate coherently in those languages.

For companies developing custom applications with artificial intelligence, this conclusion has immediate practical implications. For example, when implementing AI agents that serve users in multiple regions, the prompting strategy should not be limited to translating the input; it is more effective to design reasoning chains in English and then output the response in the user's language. This technique also allows better leveraging of larger models, where methods based on verbalizing uncertainty (such as asking the model to express its confidence in natural language) outperform traditional probabilistic approaches, while in small models, open-box methods based on token probabilities remain more effective.

From a software engineering perspective, integrating these uncertainty estimation mechanisms requires robust infrastructure. Many organizations opt for AWS and Azure cloud services to scale inference processing and manage dynamic abstention thresholds. Additionally, combining them with business intelligence services and Power BI tools enables real-time monitoring of response reliability and adjustment of confidence criteria according to the business context. It is not just about technical accuracy: cybersecurity also comes into play, as a model that does not know when to abstain can leak sensitive information or generate dangerous hallucinations in regulated environments.

At Q2BSTUDIO, as a software and technology development company, we address these challenges by offering AI for businesses that natively integrates multilingual uncertainty estimation. Our team implements custom software solutions that leverage the latest advances in AI agents, ensuring that each interaction is backed by statistical quality control. Whether for customer service chatbots, sales assistants, or document analysis systems, the ability to measure when a model should stay silent is as valuable as its ability to speak. Because, in the end, true artificial intelligence is not just about generating responses, but about knowing when it is better not to.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.