{"article":{"slug":"the-agent-said-it-was-done-the-database-disagreed","title":"The Agent Said It Was Done. The Database Disagreed.","subtitle":null,"summary":"Microsoft’s ThinkingBox grades AI agents on backend records they leave behind—not chat claims—running isolated MCP tool sessions and scoring terminal state consistency across repeated trials; now available on Hugging Face.","content_type":"announcement","language":"en","canonical_url":"https://huggingface.co/blog/microsoft/thinkingbox","author":{"name":"Tuhin Kundu","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Microsoft","url":"https://huggingface.co/microsoft","listing_slug":null,"listing":null},"topics":[{"name":"AI Agents","slug":"ai-agents","url":"https://listedarticles.com/topics/ai-agents"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"Developer Tools","slug":"developer-tools","url":"https://listedarticles.com/topics/developer-tools"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3044,"reading_minutes":13,"published_at":"2026-10-03T00:00:00.000Z","added_at":"2026-10-04T11:14:36.133Z","updated_at":"2026-10-04T11:14:36.133Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/the-agent-said-it-was-done-the-database-disagreed","markdown_url":"https://listedarticles.com/articles/the-agent-said-it-was-done-the-database-disagreed.md","example":false,"citation":"Tuhin Kundu, Microsoft. \"The Agent Said It Was Done. The Database Disagreed..\" 3 Oct 2026. https://huggingface.co/blog/microsoft/thinkingbox (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://huggingface.co/blog/microsoft/thinkingbox"},"body_markdown":"# The Agent Said It Was Done. The Database Disagreed.\n\n*Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row. It is now available through Hugging Face.*\n\n*Figure 1: ThinkingBox runs an agent against isolated MCP tool sessions, then grades the terminal backend state and side effects it leaves behind. From our ThinkingBox paper.*\n\n#### This is a joint blog by Microsoft and Hugging Face, special thanks to Tommy Guy (founder at Enderis AI, previously Microsoft), Sergio Paniego from Hugging Face and our former interns Zhuochun Li (University of Pittsburgh), Ali Keramati (UC Irvine), Youngmin Ko (Northwestern) for co-authoring/reviewing efforts.\n\nA customer writes in. Her $745 kitchen appliance has been stuck in a courier \"exception\" at a Nashville distribution center, fifteen days past its estimated delivery date.\n\nThe AI agent does careful work. Nine tool calls: it pulls the order, checks tracking, looks up her customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and reads the policy correctly; her account segment genuinely does not qualify for late-delivery compensation.\n\nThen it closes the ticket as **resolved** and replies “ *Since your query is resolved, is there anything I may assist you with?* ”\n\nTwo things are wrong. The carrier exception is still open, so the required end state was **on hold**, pending resolution. And the customer never got a real answer to what she actually asked.\n\nAn AI grader checking tool calls would see nine well-formed ones. The grader checking whether the agent wrote to the database would see that too. The database is what **disagrees**.\n\nThat gap is what ThinkingBox measures. Across 507 stateful business workflows, each run 20 times against various LLM models, it grades agents on terminal backend state and side effects. This post covers what we found, what consistency costs, and how to run the benchmark yourself through OpenEnv.\n\n**You can run this one yourself:** the example above is adapted from a benchmark task sandbox_external_retail_group1.py:test_case_ST003_006, and the executable check that fails is a single field: the ticket's status is solved where the required end state is hold. The full trace is in  Appendix D.4, Case 3 of our paper.\n\n**Contents**\n\n*Want to try it before reading the results? Skip to section Run it yourself.*\n\n## A tool call is not an outcome\n\nFinal responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. Only the records it leaves behind settle the question.\n\nThe gap is substantial. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those state-check findings overlap.\n\nA trajectory is a claim. Database state is the evidence. Repetition is the trust test.\n\n## One success is not reliability\n\nAn agent that processes a refund correctly once and mishandles it the next four times is not a working refund agent. So every task runs **20 independent times**, each from an identical clean backend, and we report three different things:\n\n*Table 1: The three numbers we report, and the question each one answers.*\n\n| Metric | What it measures | What it answers | \n|---|---|---|\n| pass@1 | Share of all attempts that succeeded | How does it usually do? | \n| pass@20 | Share of tasks solved **at least once** in 20 tries | Can it *ever* do this? Breadth. | \n| Observed 20/20 | Tasks that actually passed **all** 20 recorded attempts | Can it *always* be correct? | \n\nWe use **observed 20/20** in this blog post as the literal count of how many of the 507 tasks passed 20 out of 20. No estimator, no smoothing. \n\nStarting with the familiar view. The table below reports pass@1, the single-attempt score estimate, broken out by domain. This is the number most leaderboards publish, and on its own it reads like an ordinary capability ranking.\n\n*Table 2: ThinkingBox-Bench pass@1 (%) by domain. Each model is evaluated on every task for 20 repeated trials. Bold marks the group leader; underline marks the runner-up. The standard errors for the single attempt score estimates are provided in Table 4 in our ThinkingBox paper.*\n\n| Model | Retail (98) | Auto insurance (100) | Travel (104) | Neobank (104) | Consulting (101) | Overall, task-weighted (507) | \n|---|---|---|---|---|---|---|\n| *Proprietary models* |  |  |  |  |  |  | \n| Claude Opus 5.5 | **80.97** | **68.40** | 54.28 | **71.25** | __61.58__ | **67.16** | \n| Claude Opus 5 | __80.71__ | __65.80__ | 49.95 | __70.62__ | **66.19** | __66.50__ | \n| GPT-5.4 | 76.33 | 62.65 | **68.12** | 65.34 | 54.60 | 65.36 | \n| GPT-5.6 Sol | 67.65 | 65.30 | __60.34__ | 59.09 | 57.52 | 61.91 | \n| Claude Sonnet 4.6 | 72.35 | 54.40 | 58.94 | 56.39 | 54.31 | 59.19 | \n| GPT-6 Astra | 71.73 | 46.55 | 55.87 | 60.87 | 56.83 | 58.31 | \n| GPT-5.2 | 70.20 | 22.40 | 53.70 | 51.15 | 34.06 | 46.28 | \n| Claude Opus 4.6 | 68.62 | 8.30 | 21.11 | 35.67 | 27.82 | 32.09 | \n| o3-pro | 37.70 | 2.95 | 17.31 | 24.28 | 14.60 | 19.31 | \n| Grok-4.3 | 43.93 | 2.60 | 15.14 | 1.78 | 9.55 | 14.38 | \n| *Open-weight models* |  |  |  |  |  |  | \n| Kimi-K3 | **82.24** | **50.80** | **61.83** | 41.35 | **51.63** | **57.37** | \n| Qwen3.8-27B | 64.03 | __47.85__ | __53.41__ | **47.88** | __45.69__ | __51.70__ | \n| DeepSeek-V4-Pro | __68.21__ | 29.65 | 43.13 | __44.86__ | 31.04 | 43.26 | \n| Kimi-K2.6 | 53.72 | 24.50 | 39.52 | 33.65 | 37.33 | 37.66 | \n| GLM-5.1 | 58.67 | 25.70 | 35.43 | 13.27 | 34.06 | 33.19 | \n| Qwen3.6-27B | 43.11 | 29.00 | 46.39 | 27.84 | 18.37 | 32.94 | \n| Qwen3.5-9B | 19.90 | 0.70 | 4.71 | 1.15 | 2.33 | 5.65 | \n| Mistral-Large-3 | 11.28 | 1.30 | 8.99 | 1.15 | 0.74 | 4.66 | \n\nClaude Opus 5.5 **leads** overall at 67.16%, two-thirds of a point above Claude Opus 5. Kimi-K3 is the strongest **open-weights model**, within a point of GPT-6-Astra. **Domain matters just as much**: Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance.\n\nOne good run tells you a model can do the work. It does not tell you whether it will do it again. So run every task 20 times and ask how much of that score survives.\n\n*Figure 2: How much of each model's single-attempt score survives 20 repeats.*\n\n**Only three** hold on to most of their pass@1 scores: GPT-6 Astra retains 78% of its single-attempt rate, and Claude Opus 5.5 and Claude Opus 5 each retain 71%. At the other end, GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro each keep about 8%.\n\nThe gap between what a model **can do once** and what it does **every time** is the whole story.\n\n## Can you depend on the model behind your agent?\n\n*Figure 3: Breadth and consistency pull apart. Twelve of the eighteen models are shown; six below 33% pass@1 are omitted for legibility.*\n\n**Kimi-K3 has the broadest coverage of any model we tested.** It solves **93.89%** of the benchmark at least once: 476 of 507 tasks. Only 31 tasks defeat it entirely, the lowest count in the field. On retail workflows it leads outright at 82.24% pass@1, ahead of every proprietary model.\n\n**Kimi-K3 is also among the least consistent.** Just 68 of 507 tasks, 13.41%, succeed in all 20 attempts.\n\nClaude Opus 5 inverts this. It solves fewer tasks at least once (79.09%; 106 defeat it entirely) but completes **47.53%** of the benchmark on every single attempt.\n\n**A newer model does not fix this.** Claude Opus 5.5 scores higher than Claude Opus 5 on every-attempt average, 67.16% against 66.50%, and solves more tasks at least once. It passes exactly the same number of tasks on all 20 attempts: 241. Half a point of headline accuracy bought no additional dependability at all.\n\n- Kimi-K3 solves **75 more tasks at least once** than Opus 5.\n- Opus 5 solves **173 more tasks consistently** than Kimi-K3.\n\nIf you are choosing a model for work that touches real records, pass@20 is the wrong column to look at.\n\n## What consistency costs\n\nCapability comparisons usually stop at the score. For anyone deploying, the relevant question is what a successful unit of work costs. We measure that as cost per successful task attempt. We say task attempt because every benchmark task is run repeatedly and cost is incurred per attempt, so pass@1 is the matching quality denominator.\n\nWe took each model's recorded token usage from its full 507 × 20 campaign and priced it at undiscounted list rates available on OpenRouter<sup>+</sup>, reversing promotional discounts and excluding endpoints that declare quantization. Input, output and cache rates all come from one provider endpoint per model.\n\nThen we divided one run's cost by the number of attempts that **succeeded**:\n\n**Cost per successful task attempt = estimated cost for 507 attempts, one per task ÷ (507 × pass@1)**\n\nThis is a comparative efficiency index, not an invoice, and not the price of serving one production request. It also prices single successes, not consistency. We price consistency next.\n\n**Example:** GPT-5.4 costs $43.49 for 507 attempts (one attempt per task) and has 65.36% pass@1, so $43.49 ÷ (507 × 0.6536) = $0.131 per successful task attempt.\n\n### Pareto cost frontier\n\nA model is on the frontier if no other model is both *no more expensive* **and** *at least as accurate*. Three models qualify; every other model is dominated on at least one axis.\n\n*Figure 4: Cost per successful task attempt against pass@1. Ringed dots are pareto cost frontier models.*\n\n**The frontier has three steps.** GPT-5.6 Sol has the **lowest cost per success** at $0.127; GPT-5.4 raises pass@1 by 3.45 percentage points for $0.004 more per success; Claude Opus 5.5 adds another 1.80 points at $0.276 per success. Each of the three remains on the cost frontier line because no cheaper model matches its pass@1.\n\nClaude Opus 5 is the clearest case: at $0.475 per successful attempt and 66.50% pass@1, it is both more expensive and less accurate than Claude Opus 5.5 at $0.276 and 67.16%.\n\n### Now price consistency\n\nCost per success rewards a model that is cheap and often right. It does not reward a model that is right every time. So we also compute cost per dependable task: the cost of the full 20-run campaign divided by the number of tasks the model passed on all 20 attempts.\n\n**Cost per dependable task = estimated cost of 20 runs of 507 attempts ÷ tasks passing 20/20**\n\n**Example:** GPT-6-Astra costs 20 × $86.03 = $1,720.60 for the campaign and passes 231 tasks on every attempt, so $1,720.60 ÷ 231 = $7.45 per dependable task.\n\n*Table 3: The nine lowest costs per dependable task among models with at least one observed 20/20 task, sorted low to high. Estimated $, not actual cloud bills.*\n\n| Model | Tasks passing 20/20 | Est. cost, 20 runs | Cost per dependable task | \n|---|---|---|---|\n| GPT-5.4 | 128 (25.25%) | $869.80 | **$6.80** | \n| GPT-6 Astra | 231 (45.56%) | $1,720.60 | $7.45 | \n| Claude Opus 5.5 | 241 (47.53%) | $1,880.77 | $7.80 | \n| GPT-5.6 Sol | 82 (16.17%) | $800.00 | $9.76 | \n| Claude Opus 5 | 241 (47.53%) | $3,206.00 | $13.30 | \n| Claude Sonnet 4.6 | 102 (20.12%) | $1,587.60 | $15.56 | \n| GPT-5.2 | 44 (8.68%) | $878.00 | $19.95 | \n| Kimi-K3 | 68 (13.41%) | $1,406.40 | $20.68 | \n| Qwen3.8-27B | 38 (7.50%) | $925.80 | $24.36 | \n\n**Now rank by consistency.** GPT-5.4 is the **cheapest** at $6.80, though only 128 tasks meet the bar. GPT-6 Astra reaches 231 at $7.45, and Claude Opus 5.5 the joint-highest 241 at $7.80.\n\nNone of the three dominates the others: each additional dependable task costs more. Claude Opus 5 also passes 241, but at $13.30, so Opus 5.5 dominates it outright. GPT-5.6 Sol, the cheapest per single success at $0.127, costs $9.76 per dependable task. The cheapest way to get a right answer is not the cheapest way to get a dependable one.\n\n## Failure signatures\n\nWe assign each failed trace one deterministic diagnostic signature, and the headline is actionable: roughly four in five failures are tool handling, not reasoning. Across an ablation study in Table 5 of our paper:\n\n| Failure signature | Share of failures | \n|---|---|\n| Tool usage | 79.9% | \n| Wrong state updates | 10.3% | \n| Incomplete user resolutions | 7.0% | \n| No state-changing action | 2.9% | \n\nThese are unweighted averages of per-model shares and observable labels, not unique causal explanations.\n\nThe practical pattern is simple: agents usually get far enough to attempt the workflow, then fail to recover from tool errors, failed preconditions, or empty lookups. That is a retry and error-recovery problem before it is a model problem.\n\nDifficulty also changes by domain: across the models listed in Table 2 above, retail averages 59.52% pass@1 while auto insurance averages 33.83%.\n\n**What to do about it.** Treat the 20/20 rate as a design input, not a verdict. The same signal the benchmark grades on is available in production: check the terminal state before you commit, not the model's summary of it.\n\nClassify tool and system errors so retries target the recoverable ones. Cut the tool surface to what the workflow needs. And require human approval on the changes you cannot cheaply reverse. We have not measured the lift from any of these on this benchmark, which is exactly the kind of thing the environment now makes testable.\n\n## How it works\n\nThinkingBox is the agent sandbox, while ThinkingBox-Bench is a dataset benchmark to evaluate agents. The diagram at the top of this post shows the loop; here is what each part does.\n\n*Figure 5: The sandbox loop from panel A in Figure 1 above: isolated tool session, terminal database state, side effects, executable judges.*\n\nEach task defines a starting backend state, a user goal, the available MCP tools, the domain policy, and executable checks over the terminal state. A simulated user holds private context (a booking reference, a preference, or a date of birth) and releases it only when asked.\n\nEvery attempt gets an isolated MCP session with freshly initialized state. Two attempts of the same task never share a database row or cached tool state, which is what makes 20-trial comparison meaningful.\n\nAt the end, a side-effect extractor derives what actually changed, and deterministic judges compare it against the required end state, accepting *any* trajectory that produces the right outcome while rejecting wrong, missing or extra effects. For requirements with no clean database value (\"did the agent disclose this is not guaranteed?\"), a narrow binary rubric question handles the semantics. 477 of the 507 tasks are graded on state alone; 30 add response rubrics.\n\nThe trust boundary: the model sees tasks, dialogue and tool schemas. Golden state, assertions, grading internals and credentials stay on the evaluator side.\n\n## Run it yourself\n\nThinkingBox is now on Hugging Face, both the harness and the dataset. ThinkingBox-Bench now sits behind the OpenEnv interface, and each finished episode returns a binary pass/fail reward. The released adapter is designed for evaluation; separate, non-benchmark scenarios can use the same interface in training workflows.\n\n### Before you start\n\nTested on Linux and WSL, with Python 3.11+, uv and Docker. You also need a thinkingbox-data checkout at the pinned release and model endpoints for the agent, simulated user and judge. One endpoint can serve all three roles, which is the simplest way to start. The OpenEnv image starts *only the OpenEnv API*; everything else you run yourself.\n\n### Install\n\n```\n# 1. OpenEnv + the ThinkingBox environment\ngit clone https://github.com/huggingface/OpenEnv\ncd OpenEnv\nuv sync --project envs/thinkingbox_env --frozen\n# 2. The executable benchmark, at the pinned release\ngit clone https://github.com/microsoft/thinkingbox-data\ngit -C thinkingbox-data checkout thinkingbox-bench-v1.0\n# 3. The ThinkingBox CLI, which provides `tb`\nuv tool install \"thinkingbox @ git+https://github.com/microsoft/thinkingbox\"\n```\n### Start Typesense\n\nIn a **second terminal**, start Typesense 30.1 and wait for its health check:\n\n```\nmkdir -p .typesense-data\ndocker run --rm -d --name thinkingbox-typesense \\\n  -p 8108:8108 \\\n  -v \"$PWD/.typesense-data:/data\" \\\n  typesense/typesense:30.1 \\\n  --data-dir /data --api-key=Fake --enable-cors\nuntil curl -fsS http://127.0.0.1:8108/health; do sleep 1; done\n```\n### Start the MCP servers\n\nIn a **third terminal**, start the Session Proxy and MCP servers.\n\n```\ncd OpenEnv\ntb mcp-start --host 127.0.0.1 --port 7111 \\\n  --servers \"$PWD/thinkingbox-data/servers/servers.yaml\"\ncurl -fsS http://127.0.0.1:7111/health\n```\n### Start the OpenEnv server\n\nBack in the first terminal, start the OpenEnv server against a ThinkingBox YAML config naming your three models (config guide):\n\n```\nOPENENV_TB_CONFIG=\"$PWD/thinkingbox.yaml\" \\\nuv run --project envs/thinkingbox_env --frozen server\n```\n### Check readiness\n\nGate on readiness before running anything. It returns 503 until its observable data, configuration and Session Proxy checks pass. It cannot observe Typesense or live-probe every model endpoint, so confirm those separately:\n\n```\ncurl -sS http://127.0.0.1:8000/ready\n```\n### Score an episode\n\nNow score a real episode. example_usage.py only resets and lists tools; for agent actions, effects and assertions use the packaged evaluator:\n\n```\necho \"- sandbox_external_retail_group1.py:test_case_ST002_001\" > one_task.yaml\nuv run --project envs/thinkingbox_env thinkingbox-eval \\\n  one_task.yaml \\\n  --config \"$PWD/thinkingbox.yaml\" \\\n  --output results.jsonl \\\n  --errors-output errors.jsonl \\\n  --repeat 1 --message-timeout 1800\n```\nThe OpenEnv adapter writes operational failures to an errors sidecar so they can be rerun rather than silently mixed with model outcomes. A canonical result must resolve or explicitly account for those attempts; we counted system errors as unsuccessful trials.\n\nRuns are gated on a pinned framework commit, a pinned data release and a bundle hash, so a canonical result is verifiable rather than asserted.\n\n## Where this goes next\n\nThe useful part of this work is not our pass@1 leaderboard. It is the environment.\n\nIf you are evaluating an agent that touches real records:\n\n1. **Inspect a failure.** Find a run that terminated cleanly and still failed, and look at what actually changed in the database. It reframes what your own evals measure.\n2. **Reproduce one task** through OpenEnv with your own model.\n3. **Report a repeat metric, and define it.** Whatever k your use case justifies; say whether you are reporting best-of-k or every-of-k, and how you computed it.\n\nFurther details can be found in the following links:\n\n- **Environment:** envs/thinkingbox_env\n- **OpenEnv:** https://huggingface.co/docs/openenv/environments/thinkingbox\n- **Framework:** microsoft/thinkingbox · tutorial\n- **Benchmark:** microsoft/thinkingbox-data · v1.0 release\n- **Dataset viewer:** microsoft/ThinkingBox-Bench\n- **Paper:** arXiv:2608.19741 or HF\n- **RL training (coming soon):** microsoft/thinkingbox-training\n\nThinkingBox code is MIT-licensed; the benchmark data is CDLA-Permissive-2.0; the OpenEnv environment ships under OpenEnv's BSD-3-Clause.\n\nDisclaimer: every task in the public benchmark is a synthetic reconstruction. The workflows and policies are modeled on real AI agentic enterprise patterns; the customers are not real.\n\n*ThinkingBox and ThinkingBox-Bench is built by the Microsoft Copilot Studio team in partnership with Toloka with collaborators from the University of Pittsburgh, Northwestern University, Columbia University and UC Irvine who interned at Microsoft. Questions are welcome in the comments section below or on github*\n\n<sup>+</sup> OpenRouter cost snapshot taken on Sept 20th 2026; Opus 5.5 pricing as per the Anthropic site.","body_html":"<h1 id=\"the-agent-said-it-was-done-the-database-disagreed\">The Agent Said It Was Done. The Database Disagreed.</h1>\n<p><em>Microsoft ThinkingBox grades AI agents on the records they leave behind, not the sentences they generate, and then asks whether they can do it twenty times in a row. It is now available through Hugging Face.</em></p>\n<p><em>Figure 1: ThinkingBox runs an agent against isolated MCP tool sessions, then grades the terminal backend state and side effects it leaves behind. From our ThinkingBox paper.</em></p>\n<h4 id=\"this-is-a-joint-blog-by-microsoft-and-hugging-face-special-thank\">This is a joint blog by Microsoft and Hugging Face, special thanks to Tommy Guy (founder at Enderis AI, previously Microsoft), Sergio Paniego from Hugging Face and our former interns Zhuochun Li (University of Pittsburgh), Ali Keramati (UC Irvine), Youngmin Ko (Northwestern) for co-authoring/reviewing efforts.</h4>\n<p>A customer writes in. Her $745 kitchen appliance has been stuck in a courier &quot;exception&quot; at a Nashville distribution center, fifteen days past its estimated delivery date.</p>\n<p>The AI agent does careful work. Nine tool calls: it pulls the order, checks tracking, looks up her customer profile, searches the refund policy twice, confirms no ticket exists, opens one, documents the timeline, and reads the policy correctly; her account segment genuinely does not qualify for late-delivery compensation.</p>\n<p>Then it closes the ticket as <strong>resolved</strong> and replies “ <em>Since your query is resolved, is there anything I may assist you with?</em> ”</p>\n<p>Two things are wrong. The carrier exception is still open, so the required end state was <strong>on hold</strong>, pending resolution. And the customer never got a real answer to what she actually asked.</p>\n<p>An AI grader checking tool calls would see nine well-formed ones. The grader checking whether the agent wrote to the database would see that too. The database is what <strong>disagrees</strong>.</p>\n<p>That gap is what ThinkingBox measures. Across 507 stateful business workflows, each run 20 times against various LLM models, it grades agents on terminal backend state and side effects. This post covers what we found, what consistency costs, and how to run the benchmark yourself through OpenEnv.</p>\n<p><strong>You can run this one yourself:</strong> the example above is adapted from a benchmark task sandbox_external_retail_group1.py:test_case_ST003_006, and the executable check that fails is a single field: the ticket&#39;s status is solved where the required end state is hold. The full trace is in  Appendix D.4, Case 3 of our paper.</p>\n<p><strong>Contents</strong></p>\n<p><em>Want to try it before reading the results? Skip to section Run it yourself.</em></p>\n<h2 id=\"a-tool-call-is-not-an-outcome\">A tool call is not an outcome</h2>\n<p>Final responses and valid tool calls are only proxies. An agent can sound correct while leaving the wrong value, changing the wrong record, or creating an extra side effect. Only the records it leaves behind settle the question.</p>\n<p>The gap is substantial. In a common-set ablation covering 121,680 valid trials across 12 LLM models, 79,853 attempts failed the executable checks. Of those failures, 67.24% still terminated cleanly, invoked a state-changing tool, and reported no final tool error. Executable checks nevertheless found wrong field values in 77.61% of them, unintended extra effects in 43.30%, and missing required effects in 25.36%. Those state-check findings overlap.</p>\n<p>A trajectory is a claim. Database state is the evidence. Repetition is the trust test.</p>\n<h2 id=\"one-success-is-not-reliability\">One success is not reliability</h2>\n<p>An agent that processes a refund correctly once and mishandles it the next four times is not a working refund agent. So every task runs <strong>20 independent times</strong>, each from an identical clean backend, and we report three different things:</p>\n<p><em>Table 1: The three numbers we report, and the question each one answers.</em></p>\n<div class=\"table-wrap\"><table><thead><tr><th>Metric</th><th>What it measures</th><th>What it answers</th></tr></thead><tbody><tr><td>pass@1</td><td>Share of all attempts that succeeded</td><td>How does it usually do?</td></tr><tr><td>pass@20</td><td>Share of tasks solved <strong>at least once</strong> in 20 tries</td><td>Can it <em>ever</em> do this? Breadth.</td></tr><tr><td>Observed 20/20</td><td>Tasks that actually passed <strong>all</strong> 20 recorded attempts</td><td>Can it <em>always</em> be correct?</td></tr></tbody></table></div>\n<p>We use <strong>observed 20/20</strong> in this blog post as the literal count of how many of the 507 tasks passed 20 out of 20. No estimator, no smoothing. </p>\n<p>Starting with the familiar view. The table below reports pass@1, the single-attempt score estimate, broken out by domain. This is the number most leaderboards publish, and on its own it reads like an ordinary capability ranking.</p>\n<p><em>Table 2: ThinkingBox-Bench pass@1 (%) by domain. Each model is evaluated on every task for 20 repeated trials. Bold marks the group leader; underline marks the runner-up. The standard errors for the single attempt score estimates are provided in Table 4 in our ThinkingBox paper.</em></p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Retail (98)</th><th>Auto insurance (100)</th><th>Travel (104)</th><th>Neobank (104)</th><th>Consulting (101)</th><th>Overall, task-weighted (507)</th></tr></thead><tbody><tr><td><em>Proprietary models</em></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>Claude Opus 5.5</td><td><strong>80.97</strong></td><td><strong>68.40</strong></td><td>54.28</td><td><strong>71.25</strong></td><td><strong>61.58</strong></td><td><strong>67.16</strong></td></tr><tr><td>Claude Opus 5</td><td><strong>80.71</strong></td><td><strong>65.80</strong></td><td>49.95</td><td><strong>70.62</strong></td><td><strong>66.19</strong></td><td><strong>66.50</strong></td></tr><tr><td>GPT-5.4</td><td>76.33</td><td>62.65</td><td><strong>68.12</strong></td><td>65.34</td><td>54.60</td><td>65.36</td></tr><tr><td>GPT-5.6 Sol</td><td>67.65</td><td>65.30</td><td><strong>60.34</strong></td><td>59.09</td><td>57.52</td><td>61.91</td></tr><tr><td>Claude Sonnet 4.6</td><td>72.35</td><td>54.40</td><td>58.94</td><td>56.39</td><td>54.31</td><td>59.19</td></tr><tr><td>GPT-6 Astra</td><td>71.73</td><td>46.55</td><td>55.87</td><td>60.87</td><td>56.83</td><td>58.31</td></tr><tr><td>GPT-5.2</td><td>70.20</td><td>22.40</td><td>53.70</td><td>51.15</td><td>34.06</td><td>46.28</td></tr><tr><td>Claude Opus 4.6</td><td>68.62</td><td>8.30</td><td>21.11</td><td>35.67</td><td>27.82</td><td>32.09</td></tr><tr><td>o3-pro</td><td>37.70</td><td>2.95</td><td>17.31</td><td>24.28</td><td>14.60</td><td>19.31</td></tr><tr><td>Grok-4.3</td><td>43.93</td><td>2.60</td><td>15.14</td><td>1.78</td><td>9.55</td><td>14.38</td></tr><tr><td><em>Open-weight models</em></td><td></td><td></td><td></td><td></td><td></td><td></td></tr><tr><td>Kimi-K3</td><td><strong>82.24</strong></td><td><strong>50.80</strong></td><td><strong>61.83</strong></td><td>41.35</td><td><strong>51.63</strong></td><td><strong>57.37</strong></td></tr><tr><td>Qwen3.8-27B</td><td>64.03</td><td><strong>47.85</strong></td><td><strong>53.41</strong></td><td><strong>47.88</strong></td><td><strong>45.69</strong></td><td><strong>51.70</strong></td></tr><tr><td>DeepSeek-V4-Pro</td><td><strong>68.21</strong></td><td>29.65</td><td>43.13</td><td><strong>44.86</strong></td><td>31.04</td><td>43.26</td></tr><tr><td>Kimi-K2.6</td><td>53.72</td><td>24.50</td><td>39.52</td><td>33.65</td><td>37.33</td><td>37.66</td></tr><tr><td>GLM-5.1</td><td>58.67</td><td>25.70</td><td>35.43</td><td>13.27</td><td>34.06</td><td>33.19</td></tr><tr><td>Qwen3.6-27B</td><td>43.11</td><td>29.00</td><td>46.39</td><td>27.84</td><td>18.37</td><td>32.94</td></tr><tr><td>Qwen3.5-9B</td><td>19.90</td><td>0.70</td><td>4.71</td><td>1.15</td><td>2.33</td><td>5.65</td></tr><tr><td>Mistral-Large-3</td><td>11.28</td><td>1.30</td><td>8.99</td><td>1.15</td><td>0.74</td><td>4.66</td></tr></tbody></table></div>\n<p>Claude Opus 5.5 <strong>leads</strong> overall at 67.16%, two-thirds of a point above Claude Opus 5. Kimi-K3 is the strongest <strong>open-weights model</strong>, within a point of GPT-6-Astra. <strong>Domain matters just as much</strong>: Claude Opus 4.6 scores 68.62% on retail but 8.30% on auto insurance.</p>\n<p>One good run tells you a model can do the work. It does not tell you whether it will do it again. So run every task 20 times and ask how much of that score survives.</p>\n<p><em>Figure 2: How much of each model&#39;s single-attempt score survives 20 repeats.</em></p>\n<p><strong>Only three</strong> hold on to most of their pass@1 scores: GPT-6 Astra retains 78% of its single-attempt rate, and Claude Opus 5.5 and Claude Opus 5 each retain 71%. At the other end, GLM-5.1, Kimi-K2.6 and DeepSeek-V4-Pro each keep about 8%.</p>\n<p>The gap between what a model <strong>can do once</strong> and what it does <strong>every time</strong> is the whole story.</p>\n<h2 id=\"can-you-depend-on-the-model-behind-your-agent\">Can you depend on the model behind your agent?</h2>\n<p><em>Figure 3: Breadth and consistency pull apart. Twelve of the eighteen models are shown; six below 33% pass@1 are omitted for legibility.</em></p>\n<p><strong>Kimi-K3 has the broadest coverage of any model we tested.</strong> It solves <strong>93.89%</strong> of the benchmark at least once: 476 of 507 tasks. Only 31 tasks defeat it entirely, the lowest count in the field. On retail workflows it leads outright at 82.24% pass@1, ahead of every proprietary model.</p>\n<p><strong>Kimi-K3 is also among the least consistent.</strong> Just 68 of 507 tasks, 13.41%, succeed in all 20 attempts.</p>\n<p>Claude Opus 5 inverts this. It solves fewer tasks at least once (79.09%; 106 defeat it entirely) but completes <strong>47.53%</strong> of the benchmark on every single attempt.</p>\n<p><strong>A newer model does not fix this.</strong> Claude Opus 5.5 scores higher than Claude Opus 5 on every-attempt average, 67.16% against 66.50%, and solves more tasks at least once. It passes exactly the same number of tasks on all 20 attempts: 241. Half a point of headline accuracy bought no additional dependability at all.</p>\n<ul><li>Kimi-K3 solves <strong>75 more tasks at least once</strong> than Opus 5.</li><li>Opus 5 solves <strong>173 more tasks consistently</strong> than Kimi-K3.</li></ul>\n<p>If you are choosing a model for work that touches real records, pass@20 is the wrong column to look at.</p>\n<h2 id=\"what-consistency-costs\">What consistency costs</h2>\n<p>Capability comparisons usually stop at the score. For anyone deploying, the relevant question is what a successful unit of work costs. We measure that as cost per successful task attempt. We say task attempt because every benchmark task is run repeatedly and cost is incurred per attempt, so pass@1 is the matching quality denominator.</p>\n<p>We took each model&#39;s recorded token usage from its full 507 × 20 campaign and priced it at undiscounted list rates available on OpenRouter&lt;sup&gt;+&lt;/sup&gt;, reversing promotional discounts and excluding endpoints that declare quantization. Input, output and cache rates all come from one provider endpoint per model.</p>\n<p>Then we divided one run&#39;s cost by the number of attempts that <strong>succeeded</strong>:</p>\n<p><strong>Cost per successful task attempt = estimated cost for 507 attempts, one per task ÷ (507 × pass@1)</strong></p>\n<p>This is a comparative efficiency index, not an invoice, and not the price of serving one production request. It also prices single successes, not consistency. We price consistency next.</p>\n<p><strong>Example:</strong> GPT-5.4 costs $43.49 for 507 attempts (one attempt per task) and has 65.36% pass@1, so $43.49 ÷ (507 × 0.6536) = $0.131 per successful task attempt.</p>\n<h3 id=\"pareto-cost-frontier\">Pareto cost frontier</h3>\n<p>A model is on the frontier if no other model is both <em>no more expensive</em> <strong>and</strong> <em>at least as accurate</em>. Three models qualify; every other model is dominated on at least one axis.</p>\n<p><em>Figure 4: Cost per successful task attempt against pass@1. Ringed dots are pareto cost frontier models.</em></p>\n<p><strong>The frontier has three steps.</strong> GPT-5.6 Sol has the <strong>lowest cost per success</strong> at $0.127; GPT-5.4 raises pass@1 by 3.45 percentage points for $0.004 more per success; Claude Opus 5.5 adds another 1.80 points at $0.276 per success. Each of the three remains on the cost frontier line because no cheaper model matches its pass@1.</p>\n<p>Claude Opus 5 is the clearest case: at $0.475 per successful attempt and 66.50% pass@1, it is both more expensive and less accurate than Claude Opus 5.5 at $0.276 and 67.16%.</p>\n<h3 id=\"now-price-consistency\">Now price consistency</h3>\n<p>Cost per success rewards a model that is cheap and often right. It does not reward a model that is right every time. So we also compute cost per dependable task: the cost of the full 20-run campaign divided by the number of tasks the model passed on all 20 attempts.</p>\n<p><strong>Cost per dependable task = estimated cost of 20 runs of 507 attempts ÷ tasks passing 20/20</strong></p>\n<p><strong>Example:</strong> GPT-6-Astra costs 20 × $86.03 = $1,720.60 for the campaign and passes 231 tasks on every attempt, so $1,720.60 ÷ 231 = $7.45 per dependable task.</p>\n<p><em>Table 3: The nine lowest costs per dependable task among models with at least one observed 20/20 task, sorted low to high. Estimated $, not actual cloud bills.</em></p>\n<div class=\"table-wrap\"><table><thead><tr><th>Model</th><th>Tasks passing 20/20</th><th>Est. cost, 20 runs</th><th>Cost per dependable task</th></tr></thead><tbody><tr><td>GPT-5.4</td><td>128 (25.25%)</td><td>$869.80</td><td><strong>$6.80</strong></td></tr><tr><td>GPT-6 Astra</td><td>231 (45.56%)</td><td>$1,720.60</td><td>$7.45</td></tr><tr><td>Claude Opus 5.5</td><td>241 (47.53%)</td><td>$1,880.77</td><td>$7.80</td></tr><tr><td>GPT-5.6 Sol</td><td>82 (16.17%)</td><td>$800.00</td><td>$9.76</td></tr><tr><td>Claude Opus 5</td><td>241 (47.53%)</td><td>$3,206.00</td><td>$13.30</td></tr><tr><td>Claude Sonnet 4.6</td><td>102 (20.12%)</td><td>$1,587.60</td><td>$15.56</td></tr><tr><td>GPT-5.2</td><td>44 (8.68%)</td><td>$878.00</td><td>$19.95</td></tr><tr><td>Kimi-K3</td><td>68 (13.41%)</td><td>$1,406.40</td><td>$20.68</td></tr><tr><td>Qwen3.8-27B</td><td>38 (7.50%)</td><td>$925.80</td><td>$24.36</td></tr></tbody></table></div>\n<p><strong>Now rank by consistency.</strong> GPT-5.4 is the <strong>cheapest</strong> at $6.80, though only 128 tasks meet the bar. GPT-6 Astra reaches 231 at $7.45, and Claude Opus 5.5 the joint-highest 241 at $7.80.</p>\n<p>None of the three dominates the others: each additional dependable task costs more. Claude Opus 5 also passes 241, but at $13.30, so Opus 5.5 dominates it outright. GPT-5.6 Sol, the cheapest per single success at $0.127, costs $9.76 per dependable task. The cheapest way to get a right answer is not the cheapest way to get a dependable one.</p>\n<h2 id=\"failure-signatures\">Failure signatures</h2>\n<p>We assign each failed trace one deterministic diagnostic signature, and the headline is actionable: roughly four in five failures are tool handling, not reasoning. Across an ablation study in Table 5 of our paper:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>Failure signature</th><th>Share of failures</th></tr></thead><tbody><tr><td>Tool usage</td><td>79.9%</td></tr><tr><td>Wrong state updates</td><td>10.3%</td></tr><tr><td>Incomplete user resolutions</td><td>7.0%</td></tr><tr><td>No state-changing action</td><td>2.9%</td></tr></tbody></table></div>\n<p>These are unweighted averages of per-model shares and observable labels, not unique causal explanations.</p>\n<p>The practical pattern is simple: agents usually get far enough to attempt the workflow, then fail to recover from tool errors, failed preconditions, or empty lookups. That is a retry and error-recovery problem before it is a model problem.</p>\n<p>Difficulty also changes by domain: across the models listed in Table 2 above, retail averages 59.52% pass@1 while auto insurance averages 33.83%.</p>\n<p><strong>What to do about it.</strong> Treat the 20/20 rate as a design input, not a verdict. The same signal the benchmark grades on is available in production: check the terminal state before you commit, not the model&#39;s summary of it.</p>\n<p>Classify tool and system errors so retries target the recoverable ones. Cut the tool surface to what the workflow needs. And require human approval on the changes you cannot cheaply reverse. We have not measured the lift from any of these on this benchmark, which is exactly the kind of thing the environment now makes testable.</p>\n<h2 id=\"how-it-works\">How it works</h2>\n<p>ThinkingBox is the agent sandbox, while ThinkingBox-Bench is a dataset benchmark to evaluate agents. The diagram at the top of this post shows the loop; here is what each part does.</p>\n<p><em>Figure 5: The sandbox loop from panel A in Figure 1 above: isolated tool session, terminal database state, side effects, executable judges.</em></p>\n<p>Each task defines a starting backend state, a user goal, the available MCP tools, the domain policy, and executable checks over the terminal state. A simulated user holds private context (a booking reference, a preference, or a date of birth) and releases it only when asked.</p>\n<p>Every attempt gets an isolated MCP session with freshly initialized state. Two attempts of the same task never share a database row or cached tool state, which is what makes 20-trial comparison meaningful.</p>\n<p>At the end, a side-effect extractor derives what actually changed, and deterministic judges compare it against the required end state, accepting <em>any</em> trajectory that produces the right outcome while rejecting wrong, missing or extra effects. For requirements with no clean database value (&quot;did the agent disclose this is not guaranteed?&quot;), a narrow binary rubric question handles the semantics. 477 of the 507 tasks are graded on state alone; 30 add response rubrics.</p>\n<p>The trust boundary: the model sees tasks, dialogue and tool schemas. Golden state, assertions, grading internals and credentials stay on the evaluator side.</p>\n<h2 id=\"run-it-yourself\">Run it yourself</h2>\n<p>ThinkingBox is now on Hugging Face, both the harness and the dataset. ThinkingBox-Bench now sits behind the OpenEnv interface, and each finished episode returns a binary pass/fail reward. The released adapter is designed for evaluation; separate, non-benchmark scenarios can use the same interface in training workflows.</p>\n<h3 id=\"before-you-start\">Before you start</h3>\n<p>Tested on Linux and WSL, with Python 3.11+, uv and Docker. You also need a thinkingbox-data checkout at the pinned release and model endpoints for the agent, simulated user and judge. One endpoint can serve all three roles, which is the simplest way to start. The OpenEnv image starts <em>only the OpenEnv API</em>; everything else you run yourself.</p>\n<h3 id=\"install\">Install</h3>\n<pre><code># 1. OpenEnv + the ThinkingBox environment\ngit clone https://github.com/huggingface/OpenEnv\ncd OpenEnv\nuv sync --project envs/thinkingbox_env --frozen\n# 2. The executable benchmark, at the pinned release\ngit clone https://github.com/microsoft/thinkingbox-data\ngit -C thinkingbox-data checkout thinkingbox-bench-v1.0\n# 3. The ThinkingBox CLI, which provides `tb`\nuv tool install &quot;thinkingbox @ git+https://github.com/microsoft/thinkingbox&quot;</code></pre>\n<h3 id=\"start-typesense\">Start Typesense</h3>\n<p>In a <strong>second terminal</strong>, start Typesense 30.1 and wait for its health check:</p>\n<pre><code>mkdir -p .typesense-data\ndocker run --rm -d --name thinkingbox-typesense \\\n  -p 8108:8108 \\\n  -v &quot;$PWD/.typesense-data:/data&quot; \\\n  typesense/typesense:30.1 \\\n  --data-dir /data --api-key=Fake --enable-cors\nuntil curl -fsS http://127.0.0.1:8108/health; do sleep 1; done</code></pre>\n<h3 id=\"start-the-mcp-servers\">Start the MCP servers</h3>\n<p>In a <strong>third terminal</strong>, start the Session Proxy and MCP servers.</p>\n<pre><code>cd OpenEnv\ntb mcp-start --host 127.0.0.1 --port 7111 \\\n  --servers &quot;$PWD/thinkingbox-data/servers/servers.yaml&quot;\ncurl -fsS http://127.0.0.1:7111/health</code></pre>\n<h3 id=\"start-the-openenv-server\">Start the OpenEnv server</h3>\n<p>Back in the first terminal, start the OpenEnv server against a ThinkingBox YAML config naming your three models (config guide):</p>\n<pre><code>OPENENV_TB_CONFIG=&quot;$PWD/thinkingbox.yaml&quot; \\\nuv run --project envs/thinkingbox_env --frozen server</code></pre>\n<h3 id=\"check-readiness\">Check readiness</h3>\n<p>Gate on readiness before running anything. It returns 503 until its observable data, configuration and Session Proxy checks pass. It cannot observe Typesense or live-probe every model endpoint, so confirm those separately:</p>\n<pre><code>curl -sS http://127.0.0.1:8000/ready</code></pre>\n<h3 id=\"score-an-episode\">Score an episode</h3>\n<p>Now score a real episode. example_usage.py only resets and lists tools; for agent actions, effects and assertions use the packaged evaluator:</p>\n<pre><code>echo &quot;- sandbox_external_retail_group1.py:test_case_ST002_001&quot; &gt; one_task.yaml\nuv run --project envs/thinkingbox_env thinkingbox-eval \\\n  one_task.yaml \\\n  --config &quot;$PWD/thinkingbox.yaml&quot; \\\n  --output results.jsonl \\\n  --errors-output errors.jsonl \\\n  --repeat 1 --message-timeout 1800</code></pre>\n<p>The OpenEnv adapter writes operational failures to an errors sidecar so they can be rerun rather than silently mixed with model outcomes. A canonical result must resolve or explicitly account for those attempts; we counted system errors as unsuccessful trials.</p>\n<p>Runs are gated on a pinned framework commit, a pinned data release and a bundle hash, so a canonical result is verifiable rather than asserted.</p>\n<h2 id=\"where-this-goes-next\">Where this goes next</h2>\n<p>The useful part of this work is not our pass@1 leaderboard. It is the environment.</p>\n<p>If you are evaluating an agent that touches real records:</p>\n<ol><li><strong>Inspect a failure.</strong> Find a run that terminated cleanly and still failed, and look at what actually changed in the database. It reframes what your own evals measure.</li><li><strong>Reproduce one task</strong> through OpenEnv with your own model.</li><li><strong>Report a repeat metric, and define it.</strong> Whatever k your use case justifies; say whether you are reporting best-of-k or every-of-k, and how you computed it.</li></ol>\n<p>Further details can be found in the following links:</p>\n<ul><li><strong>Environment:</strong> envs/thinkingbox_env</li><li><strong>OpenEnv:</strong> <a href=\"https://huggingface.co/docs/openenv/environments/thinkingbox\" rel=\"nofollow ugc noopener\">https://huggingface.co/docs/openenv/environments/thinkingbox</a></li><li><strong>Framework:</strong> microsoft/thinkingbox · tutorial</li><li><strong>Benchmark:</strong> microsoft/thinkingbox-data · v1.0 release</li><li><strong>Dataset viewer:</strong> microsoft/ThinkingBox-Bench</li><li><strong>Paper:</strong> arXiv:2608.19741 or HF</li><li><strong>RL training (coming soon):</strong> microsoft/thinkingbox-training</li></ul>\n<p>ThinkingBox code is MIT-licensed; the benchmark data is CDLA-Permissive-2.0; the OpenEnv environment ships under OpenEnv&#39;s BSD-3-Clause.</p>\n<p>Disclaimer: every task in the public benchmark is a synthetic reconstruction. The workflows and policies are modeled on real AI agentic enterprise patterns; the customers are not real.</p>\n<p><em>ThinkingBox and ThinkingBox-Bench is built by the Microsoft Copilot Studio team in partnership with Toloka with collaborators from the University of Pittsburgh, Northwestern University, Columbia University and UC Irvine who interned at Microsoft. Questions are welcome in the comments section below or on github</em></p>\n<p>&lt;sup&gt;+&lt;/sup&gt; OpenRouter cost snapshot taken on Sept 20th 2026; Opus 5.5 pricing as per the Anthropic site.</p>","headings":[{"level":1,"text":"The Agent Said It Was Done. The Database Disagreed.","id":"the-agent-said-it-was-done-the-database-disagreed"},{"level":2,"text":"A tool call is not an outcome","id":"a-tool-call-is-not-an-outcome"},{"level":2,"text":"One success is not reliability","id":"one-success-is-not-reliability"},{"level":2,"text":"Can you depend on the model behind your agent?","id":"can-you-depend-on-the-model-behind-your-agent"},{"level":2,"text":"What consistency costs","id":"what-consistency-costs"},{"level":3,"text":"Pareto cost frontier","id":"pareto-cost-frontier"},{"level":3,"text":"Now price consistency","id":"now-price-consistency"},{"level":2,"text":"Failure signatures","id":"failure-signatures"},{"level":2,"text":"How it works","id":"how-it-works"},{"level":2,"text":"Run it yourself","id":"run-it-yourself"},{"level":3,"text":"Before you start","id":"before-you-start"},{"level":3,"text":"Install","id":"install"},{"level":3,"text":"Start Typesense","id":"start-typesense"},{"level":3,"text":"Start the MCP servers","id":"start-the-mcp-servers"},{"level":3,"text":"Start the OpenEnv server","id":"start-the-openenv-server"},{"level":3,"text":"Check readiness","id":"check-readiness"},{"level":3,"text":"Score an episode","id":"score-an-episode"},{"level":2,"text":"Where this goes next","id":"where-this-goes-next"}]}}