In a 24,000-episode simulation reported this September (arXiv:2609.30662), an agent governed by a separate executive layer used 36.4% fewer tokens than the same architecture under local control without the executive layer, at essentially identical success rates. If your autonomous agent keeps retrying a failed step and burning budget, the evidence points to your orchestration layer, not the model, as the thing to fix.
The symptom: spend that scales with failure, not task length
A retry loop has a distinctive cost signature. Two tasks of identical length produce wildly different bills depending on how many times the agent attempts a step that cannot succeed. Long-horizon tasks do not inherently cost more; long-horizon tasks with unbounded failure recovery do. Each repeated attempt re-reads context, re-plans, and re-issues tool calls, so token spend tracks the count of failed attempts rather than the size of the problem.
This reframes the economics. If persistence is the driver of cost, then upgrading the model or extending the context window treats the wrong variable. A smarter model that still decides, on its own, whether to retry a failed step for the fortieth time has the same structural defect as a weaker one. The question shifts from “is the model capable enough” to “who has the authority to stop, and under what contract.”
Why it happens: the dominant loop has no retry contract
The retry loop is not an exotic edge case. A 2026 survey of 70 open-source LLM agent projects (arXiv:2604.11378) found that 60% (42 of 70) adopt the Agent Loop pattern, in which the model observes, thinks, and acts in a single self-conditioned cycle. The same paper identifies the structural defect directly: “failure recovery has no bounded semantics. When a step fails, the LLM autonomously decides whether to retry, skip, or replan, with no explicit contract specifying which recovery actions are available for which failure types, and no bound on how many attempts may be made.”
Read that carefully. In the majority pattern, there is no attempt cap, no failure taxonomy, and no separate decision-maker for replan-versus-stop. Retry loops are not a model bug to be patched with better weights; they are the default behavior of an architecture that never specified what happens after failure.
A second preprint, “LLM Parkinsonism”, locates the problem more precisely. Its authors argue that persistence after failure is “not explained by autoregressive next-token prediction alone, but more directly by concentrating proposal generation, scope interpretation, progress assessment, and stopping authority within the same self-conditioned loop.” The component that proposes an action is also the component that judges whether it worked and whether to try again. That is a conflict of interest baked into the control flow, and it parallels what we covered in why production agents fail silently: systems that trust the agent’s own report of progress instead of an independent check keep missing failures that a separate verifier catches.
The Parkinsonism paper names the resulting behavior “token-inefficient persistence”: continued action after the task-level value of further attempts has collapsed. The name is provocative; treat it as the authors’ framing rather than an established diagnosis. It is a single 2026 preprint, and everything quantitative below is a reported claim from that preprint’s simulations, not a replicated result.
Diagnosis checklist: match the behavior to the missing control
Before reaching for any fix, identify which persistence behavior you are actually seeing. Each maps to a different missing control, and applying the wrong one wastes effort.
| Observed behavior | What is missing | Control to add | Evidence and tradeoff |
|---|---|---|---|
| Same failed tool call retried indefinitely | Attempt bound per failure type | Explicit recovery contract: which failures may be retried, how many times, who escalates | The Agent Loop survey identifies unbounded recovery semantics as the structural cause (arXiv:2604.11378); cheapest fix, no architectural change |
| Agent picks a reasonable action but never reassesses the plan | Progress assessment outside the generation loop | Candidate-set local control that evaluates alternatives before committing | In the GEC simulations, local control alone moved hard-goal success from 67.42% to 96.53% (arXiv:2609.30662) |
| Agent meets local step goals while the overall task drifts or burns budget | Project-level state: scope, evidence, resource use, stopping | Global executive layer separated from action generation | Reported 36.4% mean-token reduction at equal success, simulation-only (arXiv:2609.30662) |
| Agent “knows” a step failed but misidentifies which one | Reliable step-level failure attribution | External replay or tracing, not the model’s self-report | Best LLM-judge step-level attribution is only ~14% accurate on Who&When (arXiv:2606.08275) |
| Agent repeats an unauthorized action across sessions | Memory integrity, not loop control | Audit of long-term memory writes; treat as a security incident | Persistent behavior can be an injected payload, not a control bug (arXiv:2602.15654) |
Two rows deserve emphasis. The fourth row kills a comforting assumption: “the model will notice it is failing” is not supported by evidence. On the Who&When benchmark, the best step-level attribution accuracy for LLM-judge baselines is about 14%, according to the Causal Agent Replay preprint. Agents, and the LLM judges asked to grade them, cannot reliably identify which step failed. Any control design that depends on the model correctly self-diagnosing its failed step inherits that error rate. This is also why agents derailing mid-run is so hard to catch from per-step logs: the error originates early and the trajectory compounds it while nothing in the output looks broken yet.
Three levels of control, in order of cost
Bounded recovery contracts are the floor. The Agent Loop survey’s critique implies the minimal fix: declare, per failure type, which recovery actions exist and how many attempts are permitted, with a defined escalation to replan or stop. This requires no new model calls and no architecture change. It will not catch drift or bad plan selection, but it converts unbounded retry spend into bounded retry spend, which is where the token economics live.
Candidate-set local control evaluates a set of candidate actions before committing, rather than acting on the first proposal. In the GEC preprint’s matched-candidate benchmark of 24,000 episodes under a 40,000-token ceiling (arXiv:2609.30662), this is where nearly all the success improvement came from: 67.42% hard-goal success for a first-candidate baseline and 96.53% for the candidate-set local control. That gap, almost 29 points, is the argument that action selection, not raw model capability, is the binding constraint in these simulations.
Global executive control (GEC v0.2) adds the layer the Parkinsonism paper actually proposes: an uncertainty-aware governance process that holds project-level state (scope, evidence, resource use, and stopping) separate from action generation. Its reported contribution is efficiency, not success: 96.57% versus 96.53% is a rounding-level difference, while mean token use dropped from 19,782 to 12,574 (36.4%) and restricted mean tokens to completion at the ceiling fell from 16,136 to 13,114 (18.7%) (arXiv:2609.30662). The authors also report eliminating measured pre-completion drift and sharply reducing gross complexity, with favorable overhead sensitivity through an additional 500 synthetic governance tokens per cycle.
Independent evidence points the same direction. A separate line of work on a Structured Cognitive Loop, which interposes symbolic control around neural reasoning, reports zero policy violations, prevention of redundant tool calls, and complete decision traceability in multi-step conditional reasoning experiments. Different mechanism, same conclusion: a structured control layer outside the raw generation loop prevents the repeated-action symptom. That report is also preprint-stage and self-evaluated, so it supports the direction of the recommendation rather than proving it.
What the simulations show, and what they cannot
Every number attributed to GEC above comes from mechanistic simulations, and the authors say so themselves: “These mechanistic simulations support explicit governance of scope, evidence, resource use, and stopping, while live-model validation remains necessary.” There is no independent replication. There is no production model run. The honest reading is that the preprint provides a mechanistic argument plus quantified design guidance: separating stopping authority from action generation can preserve success while cutting token spend, and the savings survive plausible governance overhead in the simulated regime.
What the simulations cannot tell you is whether a live frontier model, with its own trained-in self-correction behaviors, shows the same 36.4% saving (arXiv:2609.30662), or whether the governance layer’s prompts and state tracking cost more than the synthetic 500-token-per-cycle assumption in a real stack. It also cannot tell you whether the near-zero success delta over local control (0.04 points) generalizes; if anything, that delta warns against buying the heaviest architecture first. The staged reading of the preprint’s own results is: local candidate control buys success, global executive control buys token efficiency and auditability. Order your adoption accordingly.
Two lookalikes that will fool your loop detector
Unreliable self-diagnosis. As noted above, ~14% step-level attribution accuracy (arXiv:2606.08275) means an agent can appear to be “stuck in a loop” when the actual problem is an early, misidentified failure that later steps keep building on. Loop detection that watches for repeated identical calls catches the dumb version of persistence. It misses the version where each retry is plausibly different and equally doomed. This is the same compounding pattern behind fabricated success on long-horizon tasks, where the failure that matters is not any single step but the trajectory-level divergence between reported and actual state.
Infected persistence. The Zombie Agents preprint describes a two-phase attack: during infection, the agent reads a poisoned source while completing a benign task and writes a payload into long-term memory through its normal update process; during trigger, the payload is retrieved or carried forward and causes unauthorized tool behavior. From the outside, this looks exactly like a retry loop or a stubborn agent: the same unauthorized action, attempted persistently, across sessions. No amount of retry capping fixes it, because the behavior is the attack working as designed. If your “loop” survives a fresh context and appears across unrelated tasks, audit the memory writes before you touch the control flow.
Decision: when an executive layer beats a model upgrade
Put the evidence together and the decision tree is short.
If your agent burns tokens on repeated failed calls, start with the bounded recovery contract: enumerate failure types, cap attempts per type, and define who escalates to replan or stop. This addresses the structural default that 60% of surveyed projects ship with (arXiv:2604.11378), at near-zero cost.
If the agent selects bad actions or never reassesses, add candidate-set local control. In the only quantified comparison available, that is where the success gains were, and it is simpler than a full executive layer.
Add a GEC-style global executive layer when token economics and auditability justify it: long-horizon tasks where persistence-driven spend dominates, or workflows where you need project-level state (scope, evidence, resource use) that survives any single step. Treat the reported 36.4% token saving (arXiv:2609.30662) as simulation-backed design guidance, not a benchmark you will reproduce.
Do not spend the budget on a bigger model or a longer context window to solve this. Neither changes who holds stopping authority. And do not trust the agent to notice its own failure; the attribution evidence says it usually cannot.
The strongest limitation on all of this: the headline architecture is one preprint’s mechanistic simulations with no live-model validation and no independent replication, and the preprint evaluates no specific framework, so treat this as a control principle rather than a feature recommendation for LangGraph, AutoGen, or the OpenAI Agents SDK. The principle, separating stopping authority from action generation and bounding recovery, stands on the convergence of several 2026 preprints. The specific numbers await a live model.
Frequently Asked Questions
When should you add a global executive layer instead of upgrading the model?
Add a GEC-style global executive layer when token economics and auditability justify it: long-horizon tasks where persistence-driven spend dominates, or workflows where you need project-level state (scope, evidence, resource use) that survives any single step.

Join the discussion
Share a useful perspective or ask a question about this article.