A support agent can pass every check your team currently runs and still leave the backend wrong. Microsoft’s ThinkingBox benchmark makes this concrete: in a common-set ablation across 121,680 valid trials spanning 12 models, 79,853 attempts failed the benchmark’s executable checks, and 67.24% of those failures “terminated cleanly, invoked a state-changing tool, and reported no final tool error”. The transcript looked fine. The tool calls were well-formed. The database was wrong.
If your evaluation pipeline grades replies and call syntax, this is the number that should change how you work. The practical consequence: before shipping a workflow agent, encode each task’s required end state and forbidden side effects as executable database assertions, run each task repeatedly against isolated backends, gate on consistency across all N runs rather than on one lucky pass, and price the model per successful attempt. The rest of this article is the evidence for that recipe, plus the independent work that bounds how far you should push it.
One caveat up front, because it matters throughout: every ThinkingBox number below is author-reported by the benchmark’s authors in its own paper and announcement, and no independent replication is cited here. Treat the rankings and dollar figures as a methodology demonstration, not a settled leaderboard.
The ticket that closed itself
The benchmark’s canonical failure is worth reading closely, because it is the kind of error transcript review cannot see. In task test_case_ST003_006, a customer asks about a shipment problem. The agent works the ticket, sets its status to solved, and wraps up. According to the ThinkingBox announcement:
Two things are wrong. The carrier exception is still open, so the required end state was on hold, pending resolution. And the customer never got a real answer to what she actually asked.
A reply-quality grader sees a polite, coherent response. A tool-call grader sees a valid status update. Neither checks whether the carrier exception record still says open, which is the fact that determines what the business should do next. The agent finished the conversation. It did not finish the job.
What ThinkingBox actually measures
Thinkingbox-bench is a 507-task test set of stateful business workflows, each run 20 times per model. Grading targets terminal backend state and side effects, not the conversation. Per the announcement’s description of the harness, 477 of the 507 tasks are graded on state alone; the remaining 30 add narrow binary response rubrics for cases where the reply itself carries semantic content no database value can capture.
The mechanics are what make the numbers interpretable:
- Isolation. Each attempt runs in a fresh MCP session with freshly initialized state. In the authors’ words, “two attempts of the same task never share a database row or cached tool state” (Hugging Face blog). Without this, repetition measures caching and contamination, not reliability.
- Deterministic grading. A side-effect extractor derives what actually changed, and deterministic judges compare that diff against the required end state, “accepting any trajectory that produces the right outcome while rejecting wrong, missing or extra effects” (blog). No LLM judge decides pass/fail on the 477 state-graded tasks.
- A trust boundary. The model sees tasks, dialogue, and tool schemas; golden state, assertions, grading internals, and credentials stay on the evaluator side. Your agent cannot read the answer key.
This design mirrors a pattern appearing elsewhere in agent evaluation: TREK, a travel-planning evaluation kit, “certifies a single executable plan deterministically—no LLM judge.” The field is converging on the idea that where a check can be executed, it should be, and LLM judges should be reserved for what cannot be.
The eval recipe, distilled
Stripped of the harness specifics, the reusable procedure is:
- Write the required end state as assertions. For each workflow task, specify the terminal database fields that must hold, in the same conjunctive style as the benchmark: ticket status is
hold, carrier exception isopen, no other rows touched. - Write the forbidden effects. Enumerate what must not change. This is not optional decoration; the failure data below shows unintended extra effects in 43.30% of failures that looked clean, alongside wrong field values in 77.61% of them.
- Run each task N times against isolated backends. Fresh state per attempt, no shared rows or cached tool results between attempts.
- Gate on all-N consistency, not pass@1. A model that succeeds once in twenty tries has demonstrated the task is possible, not that it is reliable.
- Price per successful attempt, not per token or per run. Divide the full campaign cost by successes.
Steps 4 and 5 are where this departs from most current practice, and where the evidence gets interesting.
What clean-looking failures hide
Return to the ablation result. Of the 79,853 failed attempts, two thirds looked clean at the transcript level. Among those clean-looking failures, the executable checks found wrong field values in 77.61%, unintended extra effects in 43.30%, and missing required effects in 25.36% (Hugging Face blog). The percentages overlap because one attempt can commit several sins.
Read the extra-effects figure twice. In over four in ten of these failures, the agent did something it was not asked to do and nothing in the transcript flagged it. In a support context, that is the class of bug that creates incidents: for instance, a refund issued alongside a replacement, or a second ticket mutated while resolving the first. Reply review is structurally blind to it, because the reply describes intent, not effect.
Independent work supports the assertion-level granularity. FACET, a terminal-task synthesis study, found that across teacher rollouts with parseable verifier results, “89.40% of individual checks are satisfied, whereas only 20.94% of completed rollouts achieve full task success.” (FACET study) When success is conjunctive, agents routinely get most fields right and the task wrong. A grader that scores partial progress or reply plausibility will systematically overstate what your agent does.
Pass@1, pass@20, and the consistency reversal
Running each task 20 times buys you three distinct numbers, and conflating them is a common misreading of agent benchmarks. The paper states it plainly:
Across 507 tasks and 20 attempts per task, pass@20 is much higher than pass@1, while all-20 success is much lower. This reveals a discovery–reliability gap: a model may be able to discover at least one successful trajectory across multiple attempts (pass@20), yet fail to reproduce that success reliably across repeated attempts (all-20 success).
Pass@20 measures capability: can the model do this at all? All-20 success measures reliability: will it do this every time a customer hits the workflow? For a deployed agent, the second number is the one that predicts incidents.
The author-reported results show this is not a theoretical distinction. Kimi-K3 solved 75 more tasks at least once than Claude Opus 5, but Opus 5 solved 173 more tasks consistently, per the ThinkingBox announcement (unreplicated). If you selected on pass@20 or best-of-N capability, you would ship the less reliable model. That reversal is the strongest argument in the whole benchmark for paying the repetition cost.
The benchmark also separates models far more than coding benchmarks do. Among eight models evaluated on both, pass@1 spanned 19.31–66.50% (47.19 points) on Thinkingbox-bench versus 90.73–95.24% (4.51 points) on no-interpreter HumanEval+, with descriptive rank correlation ρ=0.50. A team that picks its agent model from coding leaderboard position is, on this evidence, working from nearly uninformative signal. Statefulness is where models differ.
What reliability costs
The blog prices each model’s recorded token usage from its full 507 × 20 campaign “at undiscounted list rates available on OpenRouter+, reversing promotional discounts and excluding endpoints that declare quantization” (methodology). All figures author-reported and unreplicated:
| Model | Reported result | Cost per successful attempt |
|---|---|---|
| GPT-5.6 Sol | Pareto frontier | $0.127 |
| GPT-5.4 | 65.36% pass@1 ($43.49 per 507 attempts) | $0.131 |
| GPT-5.4 vs frontier | +3.45 points pass@1 for $0.004 more than Sol | — |
| Claude Opus 5.5 | +1.80 points pass@1 over GPT-5.4 | $0.276 |
| Claude Opus 5 | Off the frontier | $0.475 |
Cost per success reframes procurement. The question stops being “what does a run cost” and becomes “what does a completed workflow cost,” which is the unit your finance team and your incident rate both care about. A model that is cheaper per token but less reliable can still end up more expensive per resolved ticket, and this table is the template for checking.
Two reasons not to treat these specific numbers as durable: they depend on point-in-time list pricing, and the paper itself records harness artifacts affecting some model comparisons and cautions that one model’s settings (Opus 4.6) were unverified. Verify against the current paper revision before citing any ranking.
Independent checks on the recipe
Three lines of outside evidence bound how far to push the recipe.
Repetition is expensive, and brute force is not the only path. Meta’s production benchmarking study reports that a central 519-question benchmark takes approximately three hours per full run, and notes that “stochasticity may also require repeated trials.” Their finding that multidimensional adaptive testing, executing 200 of 519 questions (38.5% of a full run), estimated full-benchmark scores with 1.03 percentage points of mean absolute error suggests a pragmatic middle ground: run the full N=20 campaign at release gates, and use adaptive subsets between them. I would not read that as permission to skip repetition; it is a way to spend the repetition budget where it discriminates.
Consistency results expire. A 400-run study of LLM penetration-testing agents measured run-to-run consistency directly and warned that “cloud-hosted LLM behavior is subject to provider-side updates (model weight changes, safety-training updates, API version changes) that are not visible to users. Exact replication of these results may not be possible as model versions evolve.” So all-20 success is not a certificate you earn once. It is a measurement with a shelf life, which argues for re-running consistency checks on a schedule and after any provider-side version change, not just at launch.
Terminal state is necessary, not sufficient. A survey on trustworthy agentic AI argues that outcome metrics capture what happened but not how, and that agents can “craft outputs that satisfy automated judges without improving true safety (a reflection analogue of reward hacking).” Database assertions cannot see process risk: a hallucinated justification, a policy violation en route to a correct end state, a manipulation of the rubric tasks. Where you do use an LLM judge for semantics, the survey’s guidance applies: document the judge and report its calibration against human labels.
The claim this exists to test
Watch how workflow agents are marketed. One current vendor post promises an agent that “issues the refund and sends confirmation, all within the same session”, with no backend evidence attached. That is exactly the claim class terminal-state grading exists to test. Until a vendor shows you the refund row, the ledger entry, and the absence of side effects, “resolved” is a transcript adjective.
What I would do with this
For a support or workflow agent heading toward production: adopt the five-step recipe, weight all-N consistency over pass@1 in model selection, and negotiate on cost per successful attempt. The harness itself is public on Hugging Face behind the OpenEnv interface, so the 507-task suite is runnable against your own candidates rather than taken on trust.
Hold the limits clearly, though. This establishes a grading method and author-reported failure rates on one benchmark, not that any named model is safe to deploy. The ThinkingBox numbers are unreplicated, the dollar figures are pricing snapshots, and state-only checks say nothing about process risk, dialogue quality beyond 30 rubric tasks, or effects outside instrumented backends. The strongest claim the evidence supports is directional: if your evaluation stops at the transcript, you are blind to the majority of the failures this benchmark found, and the fix is assertions on state, not better reply rubrics.
Frequently Asked Questions
How many tasks are in the ThinkingBox benchmark and how are they graded?
Thinkingbox-bench is a 507-task test set of stateful business workflows, each run 20 times per model. Grading targets terminal backend state and side effects, not the conversation. Per the announcement’s description of the harness, 477 of the 507 tasks are graded on state alone; the remaining 30 add narrow binary response rubrics for cases where the reply itself carries semantic content no database value can capture.
What is the difference between pass@20 and all-20 success?
Pass@20 measures capability: can the model do this at all? All-20 success measures reliability: will it do this every time a customer hits the workflow? For a deployed agent, the second number is the one that predicts incidents.

Join the discussion
Share a useful perspective or ask a question about this article.