Skip to content

Microsoft and Hugging Face built a benchmark that grades AI agents on the database they leave behind

ThinkingBox runs 507 stateful business workflows 20 times each and scores agents on terminal backend state, not on their tool calls or final replies.

By VibecodedThis 3 min read
Diagram of the ThinkingBox agent sandbox showing an agent running against isolated MCP tool sessions with the terminal backend state graded afterward
Microsoft / Hugging Face (ThinkingBox blog)

Microsoft and Hugging Face published a joint blog post on October 3 introducing ThinkingBox, a new way to test AI agents that ignores what the agent says and grades it on what it actually leaves behind in the database.

The opening example makes the point. A customer writes in about a $745 kitchen appliance stuck in a courier exception at a Nashville distribution center, fifteen days past its estimated delivery date. The agent does careful work: nine well-formed tool calls, pulls the order, checks tracking, reads the refund policy correctly. Then it closes the ticket as resolved and replies, "Since your query is resolved, is there anything I may assist you with?"

Two things are wrong. The carrier exception is still open, so the required end state was "on hold, pending resolution." And the customer never got a real answer to her question. An evaluator checking tool calls would see nine good ones. The database disagrees.

What ThinkingBox measures

ThinkingBox is the agent sandbox; ThinkingBox-Bench is the executable benchmark built on it. It runs 507 stateful business workflows across retail, auto insurance, travel, neobank support, and consulting IT/HR, each one repeated 20 times against 18 proprietary and open-weight models. Every attempt starts from an identical clean backend with an isolated MCP session, and the agent is graded on terminal backend state and side effects: 477 tasks are graded on state alone, 30 add response rubrics.

The post is built on the paper One Success Isn't Reliability, and the authors report three metrics for every model: pass@1 (single-attempt accuracy), pass@20 (solved at least once in 20 tries, or breadth), and observed 20/20 (tasks that actually passed all 20 attempts, or consistency).

The gap between those numbers is the whole story. In a common-set ablation covering 121,680 valid trials across 12 models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. The state checks found wrong field values in 77.61% of failures, unintended extra effects in 43.30%, and missing required effects in 25.36%. The authors estimate roughly four in five failures are tool handling, not reasoning.

The results: breadth is not dependability

Claude Opus 5.5 leads overall pass@1 at 67.16%, a hair above Claude Opus 5 at 66.50%. Kimi-K3 is the strongest open-weight model at 57.37% pass@1, and it has the broadest coverage of any model tested: it solves 93.89% of the benchmark at least once, with only 31 tasks defeating it entirely.

But Kimi-K3 is also among the least consistent. Just 68 of 507 tasks, 13.41%, succeed on all 20 attempts. Claude Opus 5 inverts this: it solves fewer tasks at least once (79.09%) but completes 47.53% of the benchmark on every single attempt.

The detail worth reading twice: Opus 5.5 scores higher than Opus 5 on every-attempt average and solves more tasks at least once, but passes exactly the same number of tasks on all 20 attempts: 241. Half a point of headline accuracy bought no additional dependability at all.

The post also prices consistency. Cost per dependable task (the full 20-run campaign divided by tasks passed on all 20 attempts) ranks GPT-5.4 cheapest at $6.80, then GPT-6 Astra at $7.45, then Claude Opus 5.5 at $7.80. Claude Opus 5 passes the same 241 tasks as Opus 5.5 but costs $13.30 per dependable task, so the newer model dominates it outright on both axes.

Why it matters for anyone shipping agents

ThinkingBox is now available through Hugging Face and the OpenEnv interface, with the dataset browsable at microsoft/ThinkingBox-Bench. The authors are blunt about the takeaway: treat the 20/20 rate as a design input, not a verdict. Check the terminal state before you commit, not the model's summary of it. Classify tool and system errors so retries target the recoverable ones, cut the tool surface to what the workflow needs, and require human approval on changes you cannot cheaply reverse.

If you are choosing a model for work that touches real records, the post's advice is direct: pass@20 is the wrong column to look at. The question that matters is how much of a model's single-attempt score survives twenty repeats, and right now that number is the one nobody publishes.