🔍 Read the full analysis: What The Database Reveals After An AI Agent Says “Done” on ThorstenMeyerAI.com
Get tech for your team delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
TL;DR
Microsoft and Hugging Face have made ThinkingBox available through Hugging Face, offering a benchmark that checks database states and side effects after AI agents perform business workflows. Its authors report frequent failures across 121,680 trials, including cases where agents used state-changing tools but did not meet the task’s executable checks.
Microsoft and Hugging Face have made ThinkingBox, a benchmark for AI agents, available through Hugging Face. It tests whether agents leave business systems in the required database state and produce the required side effects—not just whether they make valid tool calls or give plausible answers—across 507 workflows run 20 times per task.
ThinkingBox runs agents in isolated sessions using Model Context Protocol (MCP) tools, then checks the backend records and side effects against executable requirements, extending the original analysis of agents claiming success when databases disagree. The benchmark covers retail, auto insurance, travel, neobanking and consulting. The release says it can also be run through OpenEnv.
In a common-set analysis spanning 121,680 valid trials across 12 models, the authors report that 79,853 attempts failed the executable checks. Among those failed attempts, 67.24% ended without a final tool error despite the agent having invoked a state-changing tool. Failure categories overlapped: checks identified wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%.
The release reports an overall pass@1 of 67.16% for Claude Opus 5.5 and 57.37% for Kimi-K3, described as the strongest open-weight model in the AI agent development table. It says Kimi-K3 scored within one percentage point of GPT-6 Astra. These are the benchmark authors’ results in their tested setup; the supplied material does not include uncertainty estimates for the comparisons.
Why Database State Changes Agent Evaluation
For companies using agents to handle refunds, support tickets, claims or bookings, a coherent response does not prove the underlying work was completed correctly. A ticket could be closed before a case is resolved, a field could be entered incorrectly, or an unrequested change could be made. Checking the final backend state can reveal these failures even when an interaction looks successful.
Repeated trials address a different question from whether a system can succeed once: how consistently does it meet the task requirements? ThinkingBox reports pass@1, the share of individual attempts that pass; pass@20, whether a task passed at least once in 20 runs; and observed 20/20, whether it passed every recorded run. A successful result across 20 tested runs is evidence about performance in that benchmark, not a guarantee of future reliability.
The results may help developers compare systems and locate workflow weaknesses, but they do not establish how an agent will perform in a particular company’s live environment. Local records, integrations, policies and unusual requests can differ from the benchmark’s conditions.
database monitoring tools for AI workflows
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
From Tool Calls to Verified Records
The benchmark is built around a distinction between what an agent does and what its actions leave behind. A valid tool call is an intermediate step; the relevant test is whether the required records and effects exist at the end of the workflow. ThinkingBox’s executable checks assess those outcomes directly.
The release illustrates the point with a retail support task about a delayed $745 appliance order. The agent investigates the delay, opens a ticket and records a timeline. The customer does not qualify for late-delivery compensation under the policy checked by the agent, but the carrier exception is still open and the ticket is required to remain on hold. The agent instead marks it solved and replies without answering the customer’s underlying question. The authors say the check fails because the ticket status is solved rather than hold.
The release describes ThinkingBox as based on the authors’ paper, but the supplied material gives no publication date. It also does not include the complete task specifications or model configurations, limiting what can be independently assessed from the release excerpt alone.
“A tool call is not an outcome.”
— Microsoft and Hugging Face, in the ThinkingBox release
As an affiliate, we earn on qualifying purchases.
What the Benchmark Cannot Yet Establish
The reported scores and failure patterns describe the authors’ tested benchmark setup. The supplied material does not provide full uncertainty estimates for model comparisons, complete evaluation configurations, or enough detail to assess how sensitive results are to different prompts and operating conditions. The reported rankings should not be read as proof that a model will perform similarly across all business systems.
It is also unclear whether passing 20 recorded runs predicts long-term reliability. Those runs provide a bounded observation under benchmark conditions; live deployments may involve changing records, unusual requests, system integrations and policies not represented in the tasks. The source does not cite an independent replication or show that benchmark performance predicts outcomes at a specific organization.
As an affiliate, we earn on qualifying purchases.
How Teams Can Test Their Own Workflows
The release says developers and researchers can run ThinkingBox through OpenEnv using isolated MCP tool sessions. This allows them to inspect tasks and evaluate agent outcomes against executable state checks. The supplied material names no future release date or other planned milestone.
For organizations considering agents, the immediate practical step is to test the workflows they intend to automate against their own policies and records, then review both passing and failing runs. ThinkingBox’s repeated checks offer one way to examine consistency; whether its results predict performance in a particular deployment remains to be tested.
As an affiliate, we earn on qualifying purchases.
Key Questions
What does ThinkingBox measure?
It checks whether an AI agent leaves a business system in the required final state and produces the required side effects after completing a workflow.
How large is the benchmark?
The release describes 507 business workflows, with each task repeated 20 times. The reported common-set analysis covers 121,680 valid trials across 12 models.
What did the authors report about failed trials?
They report that 79,853 of the 121,680 trials failed executable checks. Among those failures, many involved wrong field values, unintended extra effects or missing required effects; those categories overlap.
Does a benchmark pass prove an agent is safe to deploy?
No. A pass shows that an attempt met the benchmark’s checks under its tested conditions. It does not establish reliability across a company’s live systems, policies or unusual cases.
Primary source: Hugging Face · via ThorstenMeyerAI.com
Halloween Picks
halloween
As an affiliate, we earn on qualifying purchases.
