STOCKTAKE: Measuring the gap between perception and action in LLM agents

STOCKTAKE benchmark separates state estimation from control in LLM agents using a fair oracle, revealing why agents fail even when they detect hidden problems.

lunes, 27 de julio de 2026 • 4 min read • Q2BSTUDIO Team

Evaluando fallos ocultos de agentes con un oráculo justo

In the world of artificial intelligence applied to business processes, LLM agents (Large Language Models) are increasingly taking on complex tasks that require sequential decisions over weeks or months. A recurring issue is that the real state of the environment —that hidden factor truly driving costs— is never directly observed. When an agent fails, we cannot tell whether it misinterpreted the available information (perception failure) or whether, even though it knew what was happening, it did not act in time (the well-known knowing-doing gap).

To address this duality, the recent academic study introducing STOCKTAKE proposes a 26-week supply-chain benchmark, modeled as a partially observable Markov decision process with six hidden factors. Its key innovation is that it allows a fair reference policy to be computed: an exact Bayes filter per factor that, fed with the same observation stream the agent receives, generates a baseline against which real performance can be measured. This yields a skill score between 0 (a symptom-blind policy) and 1 (perfect oracle), and by analyzing the agent's weekly rationales, it detects the delay in identifying hidden states and the knowing-doing rate.

Initial results, obtained with cutting-edge models such as Claude Sonnet 5, GPT-5.4, DeepSeek-V4-Pro, and Grok 4.5, reveal a paradox: all detect between 84% and 88% of hidden failures, typically within the same week of onset, yet their skill scores range from 0.62 to -0.23. That is, two of the four models end up below the symptom-blind floor, even though they identify factors almost as quickly as those that outperform it. What is going wrong?

The study identifies two faces of the same coin. On one hand, when stress persists, between 34% and 43% of weeks in which the agent correctly diagnoses the hidden state still end in stockout. This indicates under-response: the agent knows something is wrong but does not adjust its actions sufficiently. However, the paradox deepens: the two models that fall below the floor are the ones that suffer the fewest stockouts during diagnosed weeks. Their problem is not inaction but overreaction: they adjust the supply chain so aggressively that the cost of their actions outweighs the benefit of avoiding a stockout. In other words, the knowing-doing gap has two directions: not doing enough, and doing too much.

For any company looking to deploy AI agents in procurement, logistics, or financial planning, this finding is crucial. It is not enough that the agent can read the environment; it must also calibrate its responses to optimize the risk-cost trade-off. This is where the customization and monitoring capabilities of Q2BSTUDIO come into play. We are a company specialized in developing custom software that integrates artificial intelligence, process automation, and business analytics.

Our methodology builds LLM agents that not only perceive and decide but also generate traceable explanations. By combining AI with cloud AWS/Azure solutions and cybersecurity, we can deploy training and evaluation environments similar to STOCKTAKE for each business use case. Moreover, integration with BI/Power BI allows real-time visualization of both perception accuracy and action efficiency, closing the continuous improvement loop.

The STOCKTAKE benchmark reminds us that, in the realm of AI agents, state estimation accuracy is not enough. The perception-action gap can manifest in opposite ways: underaction or overaction. At Q2BSTUDIO we help companies design systems that measure both dimensions and adjust their control policies with dynamic optimization algorithms. For example, for a retail client we implemented a replenishment agent that, after detecting an abnormal demand pattern, not only triggered emergency orders but also evaluated whether the cost of those orders exceeded the possible lost profit, reducing excess inventory by 18%.

In short, the lesson from STOCKTAKE is that evaluating an LLM agent solely by its final outcome is misleading. We need tools that decompose failure into perception and action components. These tools, combined with flexible automation platforms and custom applications, allow organizations not only to deploy smarter agents but also to understand why they succeed or fail. At Q2BSTUDIO we are committed to closing that gap, offering solutions that range from AI consulting to full implementation of multi-agent systems in hybrid cloud environments.

If your company is exploring the use of autonomous agents for decision-making in supply chain, finance, or customer service, we invite you to contact us. Together we can design an evaluation framework that measures both perception and action, ensuring your agent —no matter how smart it seems— does not fall into the trap of the knowing-doing gap.

A BREAK?

Play for a moment before you go

OUR SERVICES

How we can help you

Do you have a project in mind?

Tell us your vision and we'll turn it into a software solution. Whatever the scope, we make your idea real.