All 35 model-game runs in a study of heuristic self-improving agents ended with self-scores of at least 0.70, and 15 of the resulting policies still scored below their game’s random baseline (Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents). When a vendor or internal team claims an agent “self-improved” through an automated loop, the score that backs the claim often comes from the same benchmark the agent iterated on. That score is cheap to manufacture and proves little. Before a self-improvement claim crosses a deployment gate or lands in capability-change documentation, reviewers should demand evaluation on held-out tasks the agent never saw, benchmark lineage logs, and calibrated judges. Otherwise they are not reviewing capability evidence; they are accepting the improvement loop’s own marketing.
What a self-improvement claim actually asserts
A self-improvement claim sounds simple: the agent changed, and the change made it better. Unpack the assertion and it contains three separate propositions a reviewer is implicitly signing off on. First, that a modification was promoted into the agent (a new prompt, policy, tool configuration, or fine-tune). Second, that the promotion was justified by a measurement. Third, that the measurement reflects capability on tasks the agent will actually face, not just familiarity with the tasks it practiced.
The third proposition is where claims fail. A survey of self-improvement in agentic systems, Self-Improvements in Modern Agentic Systems, states the requirement plainly: to confirm genuine capability gains rather than memorization of feedback, “evaluation should measure performance on a held-out distribution that does not overlap with the optimization data.” If the improvement loop and the evaluation draw from the same task pool, the score cannot distinguish an agent that got better from an agent that memorized its own test.
That distinction is the reviewer’s entire job. A deployment gate that accepts in-loop scores is not gating anything. It is certifying a number whose generating process included the answer key.
Self-authored verification manufactures improvement
The strongest empirical evidence comes from a study of heuristic self-improving agents, Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents. The setup let agents write their own tests and score themselves against them while improving policies for game-playing tasks. The author-reported result is stark: among the 35 model, game cells with valid self-test outputs, every single one ended with a self-score of at least 0.70. Yet 15 of those 35 completed policies scored below their game’s random reference baseline, including six policies sitting at Pong’s −21.0 floor, the worst possible outcome.
Read that as a reviewer would. Every self-improvement run reported success. Nearly half produced policies worse than random play. A governance process that accepted those self-scores would have approved fifteen regressions as improvements, and no internal signal would have flagged any of them.
The same paper proposes a remedy that matters for procurement language: a Sealed Exogenous Acceptance Loop (SEAL), which retains self-authored tests for the improvement signal but compares each candidate against the incumbent through a fixed harness-side audit the harness controls, not the agent. The mechanism is worth understanding because it generalizes. The failure was not that the agent wrote tests; it was that the agent’s tests were also the acceptance criterion. Separating the improvement signal from the acceptance decision is the structural fix, and it is a fix a reviewer can demand contractually.
The admissibility checklist
The evidence base converges on four requirements before a self-improvement score should count as capability evidence.
1. Held-out, contamination-sealed evaluation tasks. The improvement pool and the evaluation pool must be disjoint, and the held-out partition must be sealed against leakage. The agentic self-improvement survey prescribes a held-out distribution that never overlaps the optimization data. Community implementations show this is buildable: Tenacious-Bench, a synthetic benchmark for B2B sales agent alignment, was constructed with contamination-sealed held-out partitions and reports its own held-out delta (a trained judge scoring 58.7 versus a 42.3 baseline, +16.4 with p=0.003, all self-reported by the repository). That “contamination-sealed” is currently a self-attested property in such repos is itself part of the checklist: ask who verified the seal, because “we held out a test set” and “an independent party confirmed the agent never saw it” are different claims.
2. Benchmark lineage and version logging. A score is only interpretable if you know exactly what harness produced it. A pilot audit of twelve LLM agent benchmark papers, What Twelve LLM Agent Benchmark Papers Disclose About Themselves, found disclosure is systematically poor. On harness provenance, seven of eight agent benchmarks scored 0.5 on the audit’s schema and one, SWE-bench, scored 0.0; the recurring gap was that the environment image is not pinned by digest, so the environment a score was measured in cannot be reproduced exactly. On cost, none of the eight agent benchmarks report inference cost in any form, giving an author-reported cost-disclosure mean of exactly 0.00. If the field’s own benchmarks do not pin their environments, a vendor’s claim almost certainly does not either. Demand pinned harness and environment images, disclosed benchmark versions, and inference cost, or treat the score as non-reproducible.
3. Calibrated judges. Many improvement loops use an LLM as judge, which moves the trust question rather than answering it. The Evaluation Context Protocol proposes explicit floors: an LLM judge “should never be trusted until it has been calibrated against a minimum of 100 human-labeled examples” (ECP) and must reach Cohen’s kappa of at least 0.6 against human domain experts. These are protocol proposals, not audited standards, but they convert “our judge is good” into two checkable numbers. A claim that cannot produce calibration evidence for its judge has not finished its homework.
4. A representative improvement signal. Even honest evaluation cannot fix a loop trained on a skewed slice of reality. The Adaptive Data Flywheel deployment paper reports, by its own account, receiving feedback from 495 employees out of 30,000 users (flywheel paper), and its authors state this creates sampling bias that reduces the generalizability of the results. Ask whose feedback drove the improvement and whether it resembles the deployment population.
The comparison below compresses the checklist into the form a reviewer can take into a gate meeting.
| Evidence item | What to demand | What the evidence shows goes wrong without it |
|---|---|---|
| Held-out evaluation | Tasks disjoint from optimization data, contamination-sealed | Score cannot separate capability from memorized feedback (survey) |
| Verification authorship | Fixed harness-side acceptance audit, not agent-authored tests | All 35 self-verified runs scored ≥0.70 while 15 finished below random (SEAL study) |
| Benchmark lineage | Pinned environment images by digest, versions, inference cost | Zero of eight agent benchmarks report inference cost; environment images unpinned (disclosure audit) |
| Judge calibration | ≥100 human-labeled examples, Cohen’s kappa ≥0.6 vs. human experts | Uncalibrated LLM judges move the trust problem instead of solving it (ECP) |
| Feedback representativeness | Feedback sample that matches the deployment population | 495 of 30,000 users fed one loop; authors concede sampling bias (flywheel paper) |
The counterpoint worth taking seriously
The checklist above implies that reusing benchmarks is disqualifying. The anchor preprint that surfaced on 2026-09-30, Which Self-Improvements Should We Trust?, argues the opposite can sometimes hold. Its proposed framework aims at a loop-wide statistical guarantee: for a user-specified level α, with probability at least 1−α, every promoted modification across the entire self-improvement process is a genuine population improvement on the underlying task distribution. In other words, reuse with controlled error rates, rather than mandatory fresh held-out sets.
This is an author-reported, non-peer-reviewed preprint, and its method and quantitative results are unverified. But the argument changes the review conversation even if the specific framework does not survive scrutiny. A sophisticated vendor may respond to a held-out-set demand with “we control the promotion error rate instead.” That is a legitimate alternative only if its inputs are disclosed: what α was chosen, what task distribution the guarantee covers, and how the judge inside the loop was calibrated. “We control the error rate” is not the same claim as “we used held-out tasks,” and a reviewer should know which one is being offered and at what strength.
There is also evidence that improvement loops can produce real change, which sharpens rather than refutes the verification requirement. SAHOO, a recursive self-improvement alignment framework, was evaluated on 189 public-benchmark tasks (63 each from HumanEval, TruthfulQA, and GSM8K) and reports code generation and mathematical reasoning gains of 16 to 18 percent against truthfulness gains of only 3.8 percent, with truthfulness carrying a higher drift cost (a CAR of 0.60 versus around 0.67). All author-reported. The lesson for reviewers is twofold: loops are not inherently fake, and a single aggregate improvement figure can hide domain collapse. If a vendor reports one delta, ask for the per-domain breakdown.
A decision rule for the deployment gate
The practical verdict follows directly. Reject in-loop scores as capability evidence. Admit a self-improvement claim into capability-change documentation or a deployment decision only when it shows all of the following:
- Held-out, contamination-sealed evaluation tasks disjoint from the optimization data, with the seal’s verification stated (independent audit versus self-attestation).
- Benchmark lineage logs: pinned harness and environment images, disclosed benchmark versions, and inference cost.
- Judges calibrated against at least 100 human-labeled examples at Cohen’s kappa ≥ 0.6 versus human domain experts (ECP).
Alternatively, accept a statistical-guarantee claim in place of fresh held-out tasks only when α, the covered task distribution, and the judge’s calibration are themselves disclosed. Anything else goes back for held-out evaluation. Note what this rule does to incentives: a review process that accepts in-loop scores lets any team manufacture an improvement claim cheaply, which shifts the entire burden of proof onto auditors and raises the cost of legitimate capability evidence for everyone else. The gate is not just protecting one deployment; it is setting the price of the word “improved” across the organization.
What this evidence cannot settle
The limits here are substantial and should shape how confidently the checklist is applied. The anchor preprint is a single author-reported, non-peer-reviewed paper observed in the feed on 2026-09-30; its method and numbers are unverified. Nearly everything supporting it is likewise 2025, 2026 arXiv preprints, plus one community repository reporting its own results. No fetched study validates the full checklist end-to-end in a real governance setting. The kappa ≥ 0.6 and 100-label thresholds are one protocol’s proposal (ECP), not an audited standard, and no result cited here has been independently replicated. The evidence also says nothing about regulatory obligations in any particular jurisdiction.
What the evidence does settle is narrower and still useful: self-authored and in-loop verification has been observed to report success while policies fail below random baselines, benchmark lineage is rarely documented well enough to reproduce a score, and the methodological literature already treats held-out evaluation as the test that separates capability from memorization. Until independent replication arrives, that is enough to justify the gate. The claim that needs defending is not “require better evidence.” It is the vendor’s claim that the current evidence suffices.
Frequently Asked Questions
What specific calibration thresholds does the Evaluation Context Protocol propose for LLM judges?
The Evaluation Context Protocol proposes explicit floors: an LLM judge “should never be trusted until it has been calibrated against a minimum of 100 human-labeled examples” (ECP) and must reach Cohen’s kappa of at least 0.6 against human domain experts.
What did the audit of twelve LLM agent benchmark papers find regarding cost disclosure?
On cost, none of the eight agent benchmarks report inference cost in any form, giving an author-reported cost-disclosure mean of exactly 0.00.
What is the proposed remedy for self-authored verification failures in the heuristic self-improving agent study?
The same paper proposes a remedy that matters for procurement language: a Sealed Exogenous Acceptance Loop (SEAL), which retains self-authored tests for the improvement signal but compares each candidate against the incumbent through a fixed harness-side audit the harness controls, not the agent.

Join the discussion
Share a useful perspective or ask a question about this article.