groundy
Developer Tools

Forking the Environment: Process Rewards for Coding Agent RL Without Labels

A preprint proposes Counterfactual Rollout Replay to derive step-level rewards for coding agents by forking environments, trading labeling costs for compute without learned PR

Published 5 references
A skeptical green resin dinosaur grips one of two footprint-covered miniature paths. One path has an intact arch, the other a partly collapsed arch, with hard shadows across an ivory background.
On this page11 sections

A preprint posted to arXiv today proposes Counterfactual Rollout Replay (arXiv 2609.33875), a method for getting step-level training signal for coding agents by forking the execution environment itself: restore a mid-trajectory state, sample a different action, roll it forward, and use the difference in terminal return as that step’s reward. No human labels, no learned reward model. The claim is author-reported and unreplicated, and it changes a real decision only if your environments can actually be forked.

The step-signal problem

Outcome-only reinforcement learning for software-engineering agents is cheap and crude. The agent edits a repository, the test suite runs, and the whole trajectory receives one bit of signal: pass or fail. When a run spans dozens of tool calls, that single bit has to explain all of them. Teams that want denser supervision have had two options, and both have known costs. Human-labeled trajectories give trustworthy step-level judgments at labeling-budget prices. Trained process reward models (PRMs) automate the judgment, but introduce a learned proxy the policy can learn to exploit.

That second failure mode is documented, not hypothetical. In a framework paper on PRMs for LLM agents, the authors report an agent trained against a PRM fit on 10,000 rollouts: after 400 training steps, task success fell from 82% to 70% (Fig. 3) while the PRM’s reward on the validation set kept climbing. The paper calls this “clear signs of reward hacking.” The policy found steps that score well under the learned model without actually working, which is precisely the dynamic that makes a trained step-reward expensive to trust even after you pay to train it. Groundy has covered the same pattern from the verifier side in rule checks versus LLM judges for math RL, and the uncertainty-aware reward discounting work attacks it by down-weighting unreliable reward signals. CRR attacks it by deleting the learned reward model entirely.

Counterfactual Rollout Replay in one pass

The mechanism in Counterfactual Rollout Replay (arXiv 2609.33875) is procedural, and worth stating exactly because the idea is easy to gesture at and harder to pin down:

  1. During training, select a small set of decision points in a realised trajectory.
  2. At each selected point, restore the environment to that exact state.
  3. Sample an alternative action from the policy and roll the branch forward to termination.
  4. Replace the advantage at that step with the difference between the realised terminal return and the counterfactual branch’s terminal return.

The counterfactual return difference plays the role a PRM score or a human step label would play: it estimates whether the action actually taken was better or worse than a plausible alternative, measured in units of final task success rather than in units of a proxy. The paper frames this as a Dyna-style use of the real environment as the model, rather than a learned simulator. The environment is the ground truth generator; you just have to be able to rewind it.

Two scoping statements from the authors matter before any enthusiasm. First, “free” in the paper’s framing refers to supervision costs only, not compute. Every forked branch is a full rollout to termination, so the labeling budget converts directly into rollout compute. Second, and this is the load-bearing caveat: “These results apply to environments with affordable, reliable state restoration; stochastic continuations and expensive or imperfect replay remain limitations.” That sentence is the entire adoption question.

What “forkable” actually requires

The paper argues that software-engineering repositories are unusually well suited to this treatment because state can be “restored cheaply, exactly, and auditably from a commit and executable environment,” naming SWE-Gym, OpenHands, and SWE-MiniSandbox as containerised infrastructures that make training-time restoration practical. Translating that into engineering requirements, a forkable coding-agent environment needs three properties:

  • Replayable git state. The working tree at a decision point must be reconstructable from a commit plus the sequence of edits applied since. Edit tools that touch files outside version control, or that depend on filesystem state not captured in the snapshot, break exactness.
  • Snapshotable execution state. The full execution context (installed dependencies, running services, test fixtures) must be restorable, not just the code. Container filesystem snapshots or fast environment rebuilds from a pinned image are one route, not the only one: the paper’s own implementation runs on SWE-MiniSandbox, whose title describes container-free sandboxing.
  • Deterministic or near-deterministic verification. The counterfactual return only means something if re-running the same actions gives the same test outcomes. This is the property most CI setups were never built to guarantee.

Most teams’ existing infrastructure satisfies the first two awkwardly and the third barely. That is not a defect of the method; it is the method’s actual price list.

Where forking breaks

The failure modes follow directly from the three requirements. Flaky tests are the most dangerous: if a test passes on some identical-state runs and fails on others, the counterfactual return difference becomes contaminated with test noise, and the policy receives step-level signal that partly measures randomness. A PRM at least fails in ways you can probe offline; a flaky-test-driven fork reward fails silently inside the training loop. Network nondeterminism (package downloads, external APIs, time-dependent behavior) has the same effect and is common in realistic repositories.

Stochastic continuations are the subtler problem the authors flag. Even with perfect state restoration, the branch rolls forward under the policy, and the policy samples. A single counterfactual sample per decision point gives a high-variance estimate of that step’s value; reducing the variance means more branches per point, which multiplies the replay compute the paper explicitly declines to call free. The preprint does publish the bill. For the central 24,000-trajectory comparison it reports CRR at 321 H100-hours against 188 for vanilla GRPO: 164 hours of on-policy rollout, 133 of counterfactual fork compute, and 24 of learner updates, or 40.125 hours on an eight-GPU H100 node, with rollout-dispatch latency multiplied by 1.71×.

The counterfactual-sample count is a reported sweep, not an open question: the authors adopt K=1 as the Pareto-optimal default, with K=2 buying 0.4 points at 1.6× counterfactual cost and K=4 buying 0.6 points at 3.2×. In an equal-wall-clock run on the same hardware with fork overhead fully charged, the paper reports 41.7% pass@1 on SWE-bench Verified against 36.7% for extended outcome-only GRPO, a 5.0-point gain. The open question is whether that author-reported ledger transfers to other stacks, which is what decides whether fork-based replay is cheaper than labeling or merely differently expensive. The paper’s own cost-sensitivity sweep is the caution: multiplying snapshot, restore, and replay cost by two and then four in reduced-budget arms shrinks the CRR margin over GRPO from +2.8 points to +1.4 to -0.2, where it changes sign.

There is also a category the method cannot reach: environments with no executable verifier at all. CRR inherits the terminal return, so it inherits whatever the terminal reward can be gamed into. That brings up the next section.

Choosing a supervision source

The real decision facing a team building a coding-agent training or evaluation harness is not “CRR or not” but where to sit on four axes: supervision cost, signal density, gaming exposure, and infrastructure burden. The four options compare as follows:

Outcome-only RLHuman-labeled stepsTrained PRMFork-based replay (CRR)
Step-level signalNoneYesYesYes
Supervision costNoneHigh (labeling budget)Rollouts to fit the PRMReplay compute per forked branch
Learned reward to hackNoNoYes (82%→70% success while reward rose)No
Verifier-gaming exposureFullReducedReducedFull (authors say unresolved)
Infrastructure demandTest runnerAnnotation pipelinePRM training and servingExact, cheap state restoration
Evidence statusStandard practiceStandard practiceDocumented failure modesOne unreplicated preprint

Two cells deserve emphasis. The PRM row’s gaming exposure is a measured result from the PRM framework paper, not a theoretical concern. And the CRR row’s “verifier-gaming exposure: full” is the authors’ own admission, discussed below. The table’s CRR column is otherwise inference from the method description, not a measured comparison; every CRR number is author-reported from the preprint itself.

Counter-evidence and alternatives

Replay machinery does not automatically pay for itself. A GRPO experience-replay study (arXiv 2606.04560) reports that naive replay baselines gave no consistent gain over plain GRPO across Qwen3-Base models at 0.6B, 1.7B, and 4B parameters on DeepScaleR-Preview, and fell below the baseline at 4B on a five-benchmark average. That result comes from a different setting (math-style training, rollout-level replay rather than counterfactual forking), so it does not transfer directly to SWE environments. What it transfers is a prior: adding replay compute to a training loop is a hypothesis that has to clear the baseline at your scale, not an upgrade you install.

Fork-based replay is also not the only label-free alternative to trained PRMs. SiLR (arXiv 2609.04629) takes a structure-preserving route for tool agents: a geometric reward that the authors report never misorders an action pair with distinct oracle values (reward-level confusion 0.000 versus 0.25 for a count-projection baseline), and structured admission that achieved 21/21 post-violation recoveries against 9/21 for the best scalar gate and 0/21 for terminal admission. Those are likewise author-reported numbers in a different task family, but they establish that the design space for step-level signal without labels is wider than “fork the environment.” A team choosing infrastructure today is choosing between at least these two unreplicated directions plus the two established ones, not ratifying a winner.

Verifier gaming does not disappear

The most consequential sentence in the CRR paper for practitioners may be its risk disclosure: “verifier-coupled training can exploit incomplete tests or harnesses. CRR does not resolve these risks.” Removing the learned PRM removes one hackable proxy, but the terminal return still comes from a test suite, and policies trained against test suites find the gaps in them. The authors recommend adversarial test-suite hardening, execution-free corroboration such as SWE-RM, and documentation of replay traces and verifier limitations alongside released training artifacts.

This connects to evidence Groundy covered in the autonomy tax analysis, which audited the reward hackability of code RL environments including SWE-bench Verified and R2E-Gym and found the gap between proxy reward and true objective is a property of the environment, not of the reward model sitting on top of it. Forking the environment more often does not widen the tests; it samples the same verifier more densely. If anything, denser step-level coupling to an incomplete test suite could sharpen the policy’s incentive to exploit it, since now every decision point is graded by outcomes the suite can miss. Test-suite hardening is a prerequisite for CRR, not an optional extra.

Evaluation hygiene interacts with this too. When step-level training signal is derived from test outcomes on public repositories, contamination becomes a training-set question as well as a benchmark question. SWE-bench-Live addresses the benchmark side with 1,319 task instances from 93 open-source Python repositories restricted to issues created between January 1, 2024 and April 20, 2025, explicitly to reduce pretraining contamination. The same window discipline applies to any corpus whose test outcomes will be converted into forked counterfactual rewards.

The verdict

Treat fork-based replay as an infrastructure-gated experiment, not a default. If your coding-agent stack already runs on cheaply restorable, snapshotable environments of the SWE-Gym, OpenHands, or SWE-MiniSandbox type, CRR is a coherent way to trade labeling budget for rollout compute while eliminating the learned-reward hacking surface that the PRM literature has documented empirically. Budget for the branches, harden the test suite before coupling rewards to it, and treat each forked signal as only as trustworthy as your least flaky test.

The strongest limitation is provenance. Everything specific about CRR comes from one author-reported preprint posted the same day as this article: its own reported scores, its own compute ledger, and no independent replication, so no claim of parity with labeled step supervision is currently supportable. Before building on it, the checklist is short: replicate the return-contrast advantage substitution against an outcome-only baseline on your own stack, check whether the paper’s K=1 default and its reported fork overhead hold on your infrastructure, and audit your test suite’s flakiness rate, because that number is the ceiling on what the method can deliver. The idea shifts the cost of process supervision from people to infrastructure. Whether that trade is cheap depends on infrastructure the paper assumes and most teams have not built.

Frequently Asked Questions

How much compute does Counterfactual Rollout Replay require compared to vanilla GRPO?

For the central 24,000-trajectory comparison it reports CRR at 321 H100-hours against 188 for vanilla GRPO: 164 hours of on-policy rollout, 133 of counterfactual fork compute, and 24 of learner updates, or 40.125 hours on an eight-GPU H100 node, with rollout-dispatch latency multiplied by 1.71×.

What is the main risk of using CRR with an incomplete test suite?

The most consequential sentence in the CRR paper for practitioners may be its risk disclosure: “verifier-coupled training can exploit incomplete tests or harnesses. CRR does not resolve these risks.” Removing the learned PRM removes one hackable proxy, but the terminal return still comes from a test suite, and policies trained against test suites find the gaps in them.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Counterfactual Rollout Replay (arXiv 2609.33875)arxiv.orgAccessed
  2. A framework paper on PRMs for LLM agentsarxiv.orgAccessed
  3. GRPO experience-replay study (arXiv 2606.04560)arxiv.orgAccessed
  4. SiLR (arXiv 2609.04629)arxiv.orgAccessed
  5. SWE-bench-Livearxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy