What happened

Microsoft and Hugging Face have released ThinkingBox, an evaluation environment and benchmark (ThinkingBox-Bench) for AI agents that handle business tasks such as refunds, insurance claims or travel bookings. Instead of judging an agent by its final reply or by whether it made valid tool calls, ThinkingBox inspects the terminal state of the backend database and any side effects the agent left behind after it finishes. The benchmark is now available through Hugging Face's OpenEnv interface, alongside a paper describing the method and results.
The motivating example in the announcement involves a customer whose appliance is stuck with a courier. An agent makes nine tool calls, reads the refund policy correctly, and closes the support ticket as resolved. Two things are wrong: the courier exception is still open, so the ticket should have been marked "on hold", and the customer never received a real answer. A grader looking only at the tool calls or the final message would mark this a success. The database disagreed.

Why it matters

Most agent evaluations today check whether the final response sounds right or whether tool calls were well-formed. ThinkingBox's authors argue these are only proxies for what actually happened to the data. In a test covering 121,680 valid trials across 12 language models, 79,853 attempts failed the benchmark's executable checks on database state. Of those failures, about two-thirds terminated cleanly, used a state-changing tool, and reported no error message, meaning they looked successful from the outside. Checking the actual state found wrong field values in most of these failures, extra unintended effects in a large share, and missing required effects in about a quarter.
The benchmark also stresses repetition. A single successful run does not show an agent is reliable, so every one of 507 workflow tasks is run 20 times from a freshly reset backend. This produces three different numbers: how an agent does on average in one attempt, whether it can ever succeed across 20 tries, and whether it succeeds every single time.

The details

On single-attempt accuracy (pass@1) across 507 tasks in five domains (retail, auto insurance, travel, neobanking and consulting), Claude Opus 5.5 led overall at 67.16%, narrowly ahead of Claude Opus 5 at 66.50%. Among open-weight models, Kimi-K3 was strongest at 57.37%, close behind GPT-6 Astra. Performance varied sharply by domain: Claude Opus 4.6 scored 68.62% on retail but only 8.30% on auto insurance.
The gap between doing something once and doing it reliably every time turned out to be large and did not track neatly with headline scores. GPT-6 Astra retained 78% of its single-attempt rate when tasks were repeated 20 times; Claude Opus 5.5 and Claude Opus 5 each retained 71%. At the other extreme, models such as GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro kept only around 8% of their single-attempt score under repetition.
Breadth and consistency pulled in different directions. Kimi-K3 solved the most tasks at least once (476 of 507, nearly 94%), but passed all 20 attempts on just 68 tasks (13.41%). Claude Opus 5 solved fewer tasks at least once (about 79%) but passed all 20 attempts on 241 tasks, nearly half the benchmark. Notably, the newer Claude Opus 5.5 scored higher than Opus 5 on average and solved more tasks at least once, yet passed exactly the same number of tasks on all 20 attempts: 241. A half-point gain in headline accuracy bought no extra dependability.
The authors also priced reliability. Using token usage and public list prices, they computed cost per successful attempt and cost per "dependable" task (one that passed all 20 runs). GPT-5.6 Sol had the lowest cost per single success at 0.127,butcost9.76 per dependable task. GPT-5.4 was cheapest per dependable task at 6.80,followedbyGPT−6Astraat7.45 and Claude Opus 5.5 at $7.80. The cheapest way to get one right answer was not the cheapest way to get a consistently right one.
Failure analysis found that roughly four out of five failures stemmed from tool usage problems rather than reasoning errors, with smaller shares attributed to wrong state updates, incomplete resolutions for the user, and cases where no state-changing action was taken at all.
  • 507 stateful workflow tasks across retail, auto insurance, travel, neobanking and consulting, each run 20 times per model
  • 121,680 valid trials across 12 models in a common-set ablation; 79,853 failed executable state checks
  • 67.24% of those failures looked clean on the surface (no tool error, proper termination)
  • Claude Opus 5.5 leads overall pass@1 at 67.16%; Kimi-K3 leads open-weight models at 57.37%
  • About 79.9% of failures are classified as tool-handling issues rather than reasoning failures
  • Released via Hugging Face's OpenEnv interface, MIT-licensed framework and CDLA-Permissive-2.0 benchmark data

How it works

Each task defines a starting backend state, a user goal, available tools (via the Model Context Protocol, or MCP), a domain policy, and executable checks against the end state. A simulated user holds private details and reveals them only when asked. Every attempt runs in an isolated session with freshly reset state, so no two attempts share data. After an episode ends, a side-effect extractor identifies what changed, and deterministic judges compare it to the required outcome, rejecting trajectories with wrong, missing or extra effects. Most tasks (477 of 507) are graded purely on database state; the remaining 30 also use a narrow rubric question for outcomes that cannot be captured as a database value, such as whether the agent disclosed a limitation to the user.

What to watch

The authors frame ThinkingBox less as a leaderboard and more as a tool: they suggest practitioners inspect individual failures to see what actually changed in a database, reproduce tasks with their own models through OpenEnv, and report repeat-based metrics rather than single-attempt scores when evaluating agents for work that touches real records. They note they have not yet measured how much specific fixes, such as checking terminal state before committing or restricting tool access, improve results on this benchmark, calling it an open question the environment now makes testable. A dataset viewer, the benchmark code, and a training-oriented release are listed as available or forthcoming through Microsoft and Hugging Face's repositories, and the announcement states the benchmark tasks are synthetic reconstructions modeled on real enterprise patterns rather than real customer data.

Sources