groundy
Culture & Society

When AI Coding Agents Edit Their Own Rules, Who Reviews the Diff?

Self-evolving AI coding rules currently show parity with no rules, not correctness gains. Teams must treat rule diffs as code, requiring human review and audit logs to manage.

Published 7 references
A scuffed translucent green resin dinosaur holds a yellow plate above a gap in its own back. A removed green plate lies below, beside an empty ivory chair, on a warm-ivory background.
On this page13 sections

The short answer to the title’s question is: right now, probably nobody, and that is the part worth fixing before anything else. An October 2026 preprint, RuleEvolve, proposes coding agents that maintain and iteratively improve their own pool of coding rules instead of consuming a hand-written AGENTS.md file. The mechanism is plausible, the measured results are modest, and the governance gap is immediate: if the rule corpus changes without a human reviewing the diff, your team’s standards now mutate through a channel that has no owner, no audit trail, and no review thread.

The practical finding up front: treat any self-updating rule system as code under review. Keep the rule corpus in version control, require a human-approved diff for every change, verify that the agent actually obeys the rules at the behavior level, and reserve formal constraints for the invariants you cannot afford to lose. The reason for that discipline is not caution for its own sake. The best current evidence says self-evolved rules do not yet buy you correctness, so what you are actually adopting is a new maintenance and audit obligation.

Your style guide is already a prompt

The reason this question arrives now, rather than as speculation, is that team standards already reach agents as a mutable text file. The RuleEvolve paper describes the arrangement plainly: “Central to the performance of these agents is the coding rules file, often instantiated as an AGENTS.md document and prepended to the backbone model’s input context.” Those rules, the authors note, shape “the correctness and length of the generated code, as well as the associated generation cost” (RuleEvolve, arXiv:2610.00650).

That framing matters more than any single benchmark number. Once the style guide is a prompt, it stops behaving like a wiki page and starts behaving like configuration. And the field-wide norm for that configuration is static and hand-maintained: a 2025 survey of self-evolving agents observes that “most existing agent systems rely on manually crafted configurations that remain static after deployment, limiting their ability to adapt to dynamic and evolving environments.”

Hand-written rules have a second problem beyond staleness. RuleEvolve’s authors call the manual process “inherently labor-intensive and suboptimal, as human-written rules often lack task-specific alignment and can even degrade agent performance,” and their own experiments found that supplying a public hand-written rule set (sourced from the awesome-cursorrules repository) did “not consistently yield performance gains” (arXiv:2610.00650). So the case for letting rules evolve is not frivolous. The status quo costs effort and sometimes makes the agent worse. The question is what you give up when you automate the fix.

Four ways rules reach the agent, and what each one costs

The sources in hand describe four distinct delivery mechanisms, and they differ most on who writes the rule, what a reviewer can inspect, and how mature the evidence is.

MechanismWho writes the ruleWhat you reviewMeasured costEvidence status
Hand-written AGENTS.md / CLAUDE.mdA human, onceThe file itself, in PRAuthoring time; rules may degrade agent performanceIndustry norm; can underperform no rules at all (RuleEvolve)
Linters and formatters (static config)Tooling authors plus your configLint output in CINear zero at runtimeCannot express 27.9% of the Google Java standard and 59.7% of the JavaScript standard (arXiv:2602.07783)
Self-evolved rule pool (RuleEvolve)The agent, iterativelyA growing, mutable rule corpusUnknown in production; no documented adoptionOne preprint, no independent replication; the 10-seed parity study uses one backbone and one benchmark, inside a broader evaluation spanning two agent frameworks, four backbones and three benchmarks (arXiv:2610.00650)
Verified self-evolution (SEVerA)The agent, under formal constraintsConstraint satisfaction, plus the spec1.9–2.5× slowdown on two of four tasksOne preprint, four tasks, no independent replication (arXiv:2603.25111)

Two rows deserve emphasis. The linter row is not a fallback answer; the coverage gap is structural, and the next section quantifies it. The verified-evolution row is the only mechanism where a rule change is mathematically prevented from violating stated constraints, and it is also the only one with an explicit runtime price attached.

What RuleEvolve actually measured: parity, not progress

RuleEvolve’s own description is ambitious: “a self-evolving framework for coding rules” that “maintains a pool of candidate coding rules and iteratively improves them” (arXiv:2610.00650). The measured result is quieter. Across 10 independent seeds with a GPT-4.1-mini backbone on BigCodeBench, the framework reached a pass rate of 0.502±0.026, which the paper itself describes as comparable to the “No rule” baseline, where the agent receives only the task description. The clearest measured gain is conciseness: 24.3±0.4 average lines and 794.4±14 characters per generation, the most compact output among the methods evaluated (RuleEvolve).

Read that twice before adopting anything. On this evidence, self-evolved rules did not make the agent more correct than giving it no rules at all. They beat a public hand-written rule set, which is a real result, but that baseline is one the paper’s own experiments show can hurt performance. Shorter code has genuine value (review load, generation cost), yet a team justifying a self-updating rule system on correctness grounds is, as of this preprint, ahead of the evidence.

This is also single-team lab work, and the multi-seed evidence comes from one configuration: one backbone, one benchmark, ten seeds, no independent replication, and no documented production adoption. The preprint’s broader evaluation does span two agent frameworks, four backbone LLMs and three benchmarks, plus SWE-Bench Lite, where it reports the highest pass rate of 0.340; the multi-seed robustness claim, though, rests on the single configuration above. The honest summary is that the mechanism exists and the gains are unproven.

Why “we already have linters” is not an answer

A common response to the rule-governance question is that static analysis already enforces team standards, so the prompt-level rules are decoration. The linter-configuration literature disagrees with numbers. Benchmarking against two Google coding standards, an automated linter-configuration study found that “27.9% of coding standards from the Google Java coding standard and 59.7% of coding standards from the Google JavaScript coding standard are unsupported by Checkstyle and ESLint.” (arXiv:2602.07783)

That is a hard ceiling on static-config enforcement, and it explains why the prompt-level rule file exists at all. Anything linters cannot express, things like “prefer the repository’s existing error type over introducing a new one,” lives in natural language, which means it lives in the rule corpus, which means it is exactly the artifact a self-evolving system will rewrite. The enforcement stack is therefore not “linters for rules, prompts for style.” It is “linters where possible, prompt rules for the rest,” and the rest is somewhere between a quarter and three-fifths of a real standard by this measurement.

The missing review loop: transcripts and tamper-evident logs

If rules can change and linters cannot cover them, what does review look like? Two community projects sketch the answer, one at the behavior level and one at the governance level. Neither is mature, and both are worth understanding as categories rather than as endorsements.

The first category is behavioral audit. The agent-rules tool states the distinction in one line: “Linters check that your CLAUDE.md is well-formed. agent-rules checks that your agent actually followed it.” It “reads your CLAUDE.md / AGENTS.md / .cursorrules, extracts the imperative rules (the prohibitions and the mandates), then scans an agent session transcript for the places the agent broke them.” The authors describe it as zero-dependency, offline, and sending nothing off the machine; that is a self-description, not a third-party audit, but the architectural point stands regardless. File-level review tells you what the rules say. Transcript-level review tells you whether they happened.

The second category is the control plane. nexus-agents describes itself as “an autonomic control plane for your AI coding agents — Claude Code, Codex, Gemini, and OpenCode,” one that “admits work through one entry point, reviews it adversarially before it ships, records every action in a tamper-evident event log, and closes the loop by tuning where the next task goes based on what actually worked.” Its own evaluation is the model of how to caveat early tooling: “100% bug-catch on a focused synthetic dataset (n=10) and a 50% raw false-positive rate,” figures the authors themselves call “directional small-n figures, not measured rates.” Treat the numbers as a design sketch with preliminary instrumentation, and the tamper-evident event log as the feature that matters for the rules question: a rule change that lands in an append-only, tamper-evident log can at least be reconstructed and contested later. The project states the limit itself: the chain is “tamper-evident, not tamper-proof.”

There is also a measured price for making agent work reviewable at all. A paired pilot study of delegation contracts (64 runs with 192 blinded reviews) found that explicit contracts “function as a reviewability mechanism rather than a correctness mechanism on small tasks,” costing +13% agent tokens and +38% wall-clock time, with the effect larger for the weaker model tier (arXiv:2606.17099). Any mandate that rule changes arrive as reviewable diffs should be budgeted with that shape of overhead in mind: you pay tokens and latency, and what you get back is inspectability, not accuracy.

When rule changes need hard constraints

For a subset of rules, “someone should review this” is too weak. The SEVerA preprint argues that “natural-language task descriptions and testing on fixed inputs is not sufficient to completely prevent such failures. Agentic programs need formal hard constraints to ensure safety.” Its approach, verified self-evolution, lets the agent improve itself while formal behavioral specifications constrain the result, and the authors report “zero constraint violations on held-out test inputs” across four tasks “while simultaneously improving task performance over every baseline.” The cost is explicit: a 1.9–2.5× slowdown relative to the LLM baseline on two of the four tasks (arXiv:2603.25111).

The takeaway for a team lead is not “adopt formal verification.” It is that the research literature has already sorted rule enforcement into two tiers: soft review for most conventions, hard constraints for the invariants where a regression is unacceptable. A self-updating rule system that cannot distinguish those tiers will either over-constrain everything (and pay the slowdown everywhere) or under-constrain the few rules that actually needed proof.

A governance checklist before you let rules self-update

Nothing in the evidence supports banning self-evolving rules, and nothing supports trusting them by default. The checklist below is the practical verdict, with each item tied to the gap it closes.

  • Version the rule corpus like source code. Every rule change lands as a diff in the repository, with author attribution (human or agent), a timestamp, and a reason. This is what makes the change contestable in public, the way a style-guide PR already is.
  • Require a human-approved diff for every rule change. The agent may propose; a named person disposes. The reviewability pilot suggests you will pay overhead for this (+13% tokens, +38% wall-clock in that study’s contract condition) and that what you are buying is the ability to review, not better code (arXiv:2606.17099).
  • Push everything expressible into linters and formatters first. They are deterministic and cheap, but plan for the measured ceiling: 27.9% and 59.7% of two Google standards were not expressible in Checkstyle and ESLint (arXiv:2602.07783). The remainder belongs in the rule file, knowingly.
  • Audit behavior, not just the file. Add transcript-level compliance scanning of the kind agent-rules implements, so a silently ignored rule shows up as a finding rather than a surprise.
  • Log rule changes in an append-only, tamper-evident log. A hash-chained event log of the kind nexus-agents describes turns “the agent rewrote its own instructions” from an unanswerable accusation into a queryable record.
  • Pin hard invariants with formal constraints where regressions are unacceptable. SEVerA shows this is achievable at a quantified 1.9–2.5× cost on some tasks; apply it surgically (arXiv:2603.25111).
  • Do not claim correctness gains internally. RuleEvolve matched a no-rule baseline; conciseness was the clearest measured improvement (arXiv:2610.00650). Justify adoption on maintenance burden and consistency, or wait for stronger evidence.

What junior developers stop learning: a hypothesis, not a finding

One consequence sits outside every measurement cited here. The traditional review thread does double duty: it corrects the code and it teaches the author. When standards live in a rule file the agent silently maintains, a junior developer’s mistake may never generate a visible correction at all; the rule pool just absorbs the pattern. Nobody’s mistake is argued about in public, so nobody learns from the argument.

None of the sources above measures this effect. It is a hypothesis to test, and it is testable: teams adopting self-updating rules could track whether review-comment volume on style issues falls, whether onboarding time changes, and whether engineers can still articulate why a standard exists. If the hypothesis holds, the bottleneck moves to whoever curates the rule store, and that curator’s judgment becomes the de facto training program. That is a staffing decision as much as a tooling one, and it deserves to be made deliberately.

Verdict: review the rule diff like code, and name its owner

When AI coding agents edit their own rules, the review unit becomes a rules diff, and the reviewer has to be a person with a name. The mechanisms for this already exist in early form: versioned rule files, transcript-level compliance checks, tamper-evident logs, and formal constraints for the invariants that cannot bend. What does not yet exist is evidence that self-evolved rules improve correctness; the one direct measurement shows parity with using no rules at all, plus shorter output, from a single team’s preprint with no independent replication.

I would let an agent propose rule changes on work that can tolerate a bad suggestion caught in review, because the alternative, a stale hand-written file that may itself degrade the agent, has its own documented cost. I would not let it merge them unreviewed, and I would not adopt any system whose rule changes leave no queryable record. The style guide used to change only when someone argued for the change in public. If you adopt self-updating rules, your job is to keep that property on purpose, because the tooling will not preserve it by default.

Frequently Asked Questions

What is the measured performance of self-evolved rules compared to no rules?

Across 10 independent seeds with a GPT-4.1-mini backbone on BigCodeBench, the framework reached a pass rate of 0.502±0.026, which the paper itself describes as comparable to the “No rule” baseline, where the agent receives only the task description.

How much of Google’s coding standards can linters like Checkstyle and ESLint enforce?

Benchmarking against two Google coding standards, an automated linter-configuration study found that “27.9% of coding standards from the Google Java coding standard and 59.7% of coding standards from the Google JavaScript coding standard are unsupported by Checkstyle and ESLint.”

What is the cost of using formal constraints for verified self-evolution?

The cost is explicit: a 1.9–2.5× slowdown relative to the LLM baseline on two of the four tasks (arXiv:2603.25111).

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. RuleEvolvearxiv.orgAccessed
  2. 2025 survey of self-evolving agentsarxiv.orgAccessed
  3. Linter-configuration studyarxiv.orgAccessed
  4. SEVerA preprintarxiv.orgAccessed
  5. agent-rulesgithub.comAccessed
  6. nexus-agentsgithub.comAccessed
  7. Delegation contracts pilot studyarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy