ThinkingBox Benchmark Shows AI Agents Fail the Database Even When They Say the Task Is Done
Microsoft and Hugging Face published ThinkingBox on October 3, 2026, a benchmark that runs every agent task 20 times and judges by the database's final state, not by the text response. The result: a large share of agents appear to have finished the job and haven't.
The Ticket the Agent "Resolved" (But Didn't)
A customer writes in complaining that a $745 appliance has been stuck for 15 days in a carrier exception. The agent makes nine tool calls: it pulls up the order, checks the tracking, reviews the profile, reads the refund policy twice, confirms no ticket is open, opens one, documents the timeline, and applies the policy correctly (the account's tier genuinely isn't entitled to delay compensation).
Then it closes the ticket as resolved and replies:
Since your query is resolved, is there anything I may assist you with?
Agent response recorded in case ST003_006 of the ThinkingBox benchmark
Two problems: the carrier exception is still open (the required state was "pending," not "resolved"), and the customer never got a real answer to what she asked. An evaluator checking only the tool calls would see nine well-formed calls. The Hugging Face and Microsoft post uses this case, reproducible with the file sandbox_external_retail_group1.py:test_case_ST003_006, to illustrate ThinkingBox's central problem: the one who disagrees with the agent is the database.
A Successful Tool Call Is Not a Correct Outcome
ThinkingBox-Bench covers 507 stateful workflows (retail, auto insurance, travel, neobank, and consulting), each run 20 times against isolated MCP sessions. In an ablation with a common set covering 121,680 valid attempts across 12 models, 79,853 failed the executable checks on the final state.
The figure that matters most to agent builders: 67.24% of those failures ended cleanly, called a state-changing tool, and reported no tool error at all. Even so, the executable checks found wrong field values in 77.61% of them, unwanted side effects in 43.30%, and missing required effects in 25.36% (the categories overlap). In other words: the agent ended up happy, and the database recorded something else.
Three Metrics, Three Different Questions
Since each task runs 20 times from a clean backend, ThinkingBox reports three numbers, each backing a different question:
| Metric | What it measures | Question it answers |
|---|---|---|
| pass@1 | fraction of all attempts that succeeded | How does it do on average? |
| pass@20 | fraction of tasks solved at least once in 20 attempts | Can it ever do this? |
| Observed 20/20 | tasks that passed all 20 recorded attempts, no estimator | Is it always correct? |
In overall task-weighted pass@1, Claude Opus 5.5 leads with 67.16%, half a point ahead of Claude Opus 5 (66.50%); GPT-5.4 follows with 65.36%. Among open models, Kimi-K3 leads with 57.37%. Domain matters as much as model: Claude Opus 4.6 scores 68.62% in retail and only 8.30% in auto insurance.
The Split Between Breadth and Consistency
pass@1 alone hides the question that matters in production: does the same agent get it right again? Kimi-K3 has the benchmark's broadest coverage: it solves 93.89% of tasks at least once (476 of 507), with only 31 tasks defeating it completely. But it's also one of the least consistent models, with only 68 of 507 tasks (13.41%) passing all 20 attempts.
Claude Opus 5 flips the picture: it solves fewer tasks at least once (79.09%, 106 tasks never solved), but completes 47.53% of the benchmark on every single attempt. Kimi-K3 solves 75 more tasks than Opus 5 at least once; Opus 5 solves 173 more tasks consistently than Kimi-K3.
A version upgrade doesn't fix this on its own, either. Claude Opus 5.5 beats Opus 5 on average pass@1 (67.16% versus 66.50%) and solves more tasks at least once, but passes the exact same number of tasks across all 20 attempts: 241. Half a point of headline accuracy bought zero additional reliability.
A trajectory is a claim. Database state is the evidence. Repetition is the trust test.
Tuhin Kundu, author of the post, Hugging Face and Microsoft
What Reliability Costs
The team priced each model by token usage for the full campaign (507 tasks × 20 rounds) at OpenRouter's list rates, then divided cost by successful attempt. GPT-5.4 costs $43.49 for 507 attempts and has 65.36% pass@1, which works out to $0.131 per successful attempt.
On the cost-per-success frontier, only three models survive without being dominated: GPT-5.6 Sol is the cheapest ($0.127), GPT-5.4 gains 3.45 points of pass@1 for just $0.004 more, and Claude Opus 5.5 adds another 1.80 points at $0.276. Claude Opus 5, at $0.475 and 66.50% pass@1, is dominated: more expensive and less accurate than Opus 5.5.
But cost per success isn't cost per reliability. Dividing the full 20-round campaign cost by the number of tasks that passed 20 of 20 times, the ranking changes: GPT-5.4 is the cheapest at $6.80 per dependable task (but only 128 tasks, 25.25%, get there), GPT-6 Astra costs $7.45 for 231 tasks (45.56%), and Claude Opus 5.5 costs $7.80 for 241 tasks (47.53%).
Claude Opus 5 also passes 241 tasks, but at $13.30 each, dominated by Opus 5.5. GPT-5.6 Sol, the cheapest per isolated success, rises to $9.76 per dependable task: the cheapest path to a correct answer isn't the cheapest path to a reliable one.
Why It Fails: Not Reasoning, Tool Handling
Every failed trace receives a deterministic diagnostic signature, and the pattern is clearly actionable: about four out of every five failures come down to tool handling, not reasoning.
- Tool usage: 79.9% of failures
- Wrong state updates: 10.3%
- Incomplete resolutions for the user: 7.0%
- No state-changing action at all: 2.9%
The practical pattern described in the post: the agent usually gets close enough to completing the flow and then fails to recover from a tool error, a failed precondition, or an empty search. This is, first and foremost, a retry and error-recovery problem, not a model problem. Difficulty also varies by domain: retail averages 59.52% pass@1 across the listed models, versus 33.83% in auto insurance.
How It Works Under the Hood: Isolated Sessions, Deterministic Judges
Each task defines an initial backend state, a user goal, the available MCP tools, the domain policy, and executable checks on the terminal state. A simulated user holds private context (a reservation number, a date of birth) and only reveals it when asked. Each attempt gets an isolated MCP session with freshly initialized state, so that two attempts at the same task never share a database row or cache: that's what makes the 20-round comparison meaningful.
At the end, a side-effect extractor derives what actually changed, and deterministic judges compare that against the required state, accepting any trajectory that produces the right outcome and rejecting wrong, missing, or extra effects. For requirements without a clean database value ("did the agent warn that this isn't guaranteed?"), a narrow binary rubric covers the semantics; 477 of the 507 tasks are judged by state alone, 30 add response rubrics. The model under test sees tasks, dialogue, and tool schemas; golden state, assertions, judging logic, and credentials stay on the evaluator's side.
Running the Benchmark Yourself
ThinkingBox is published on Hugging Face as a harness and dataset, behind the OpenEnv interface, with each episode returning a binary pass/fail reward. The basic flow, summarized from the post itself, requires Python 3.11+, uv, and Docker:
git clone https://github.com/huggingface/OpenEnv
cd OpenEnv
uv sync --project envs/thinkingbox_env --frozen
git clone https://github.com/microsoft/thinkingbox-data
git -C thinkingbox-data checkout thinkingbox-bench-v1.0
uv tool install "thinkingbox @ git+https://github.com/microsoft/thinkingbox"Then you bring up Typesense (the index used by the MCP servers), the MCP servers themselves, and the OpenEnv server pointing to a YAML with the three model endpoints (agent, simulated user, and judge can all be the same endpoint). With everything running, you can score a real episode with thinkingbox-eval, passing a file with the task (sandbox_external_retail_group1.py:test_case_ST002_001) and getting back separate results.jsonl and errors.jsonl, so operational failures don't get mixed in with model failures. Runs are pinned to a fixed framework commit, a fixed dataset version, and a package hash: a "canonical" result is verifiable, not just claimed.
What Changes for Those Building Agents in Production
The authors' practical message is direct: treat the 20/20 rate as a design input, not a verdict. The same signal the benchmark judges by is available in production: check the terminal state before committing, not the summary the model gives of it. Classify tool and system errors so retries target only the recoverable ones, trim the tool surface to what the flow actually needs, and require human approval for changes that can't be cheaply undone. The authors themselves admit they haven't measured the actual gain from any of these practices within the benchmark, only that the environment now makes it testable.
One caveat is worth noting: all public tasks are synthetic reconstructions modeled on real enterprise-agent patterns, and the cost-per-attempt figures are a comparative index at list rates, not the actual bill for running an agent in the cloud. For whoever decides which model to put behind an agent that writes orders, refunds, or policies, ThinkingBox's lesson is that pass@1 alone measures the wrong skill: what needs to be demanded is repetition against real state, because that's what the database will record after the agent has already said it's done.
Translated from the Brazilian Portuguese original · Read the original
Claude Code refines subagents and hooks, exposing the limits of automating PR review
Three Claude Code releases published between October 1 and 3, 2026 change subagent isolation and the behavior of hooks like PreToolUse. The fixed bugs reveal how to build (and where not to blindly trust) a local code review pipeline.