groundy
Culture & Society

Cockroach Labs Spent 5 Months Treating Bugs Like Patients: What Humans Keep

Cockroach Labs reports a five-month agent triage experiment. Independent studies show agents fail at reproduction, so adopt the review gates, not the self-reported volume.

Published 4 references
A translucent green resin dinosaur presents a clear tray holding a visibly seamed ivory ring. An oversized ivory hand lifts a yellow dome above it, casting hard shadows across a warm cream background.
On this page11 sections

Cockroach Labs reports that it ran a bug-triage experiment for just under five months in which coding agents, organized like a hospital medical team, landed well over a million lines of code. Every figure in that sentence is self-reported by a single vendor running the model on its own codebase, with no independent replication and no pre-experiment baseline, so treat it as one company’s case study, not a result. Even so, the operating model is worth stealing selectively. The defensible answer for an engineering lead is: borrow the gates, not the numbers. Require agents to reproduce and diagnose before they touch code, keep precedent creation, escalation rulings, and discharge sign-off human, and measure escalation frequency and review evidence quality rather than raw fix counts. The consequence of adopting the model is that your quality-control burden shifts from writing fixes to reviewing cases, and your junior engineers lose the repair work that used to teach them, a cost the vendor itself concedes.

What actually ran

According to Cockroach Labs’ account of the experiment, the hospital pipeline ran in a mirror repository between April 21 and September 11, and: “In just under 5 months, the hospital pipeline landed well over a million lines of code, split into changes small enough that they could be confidently reviewed by agents.”

Two details in the same post keep that headline honest. First, 1,299 of the issues filed in the mirror repository, nearly half, were created by the hospital for the hospital: decomposition children and follow-ups. That is overhead work the pipeline generates for itself, so issue counts and line counts measure volume, not validated resolution. Second, the company attributes a substantial instruction burden to the design: the Discharge Nurse agent alone loads about 21,800 words, roughly 29,000 tokens, of instructions on every run before it reads a single pull request, and all fifteen agents load the base hospital-protocol skill. The company calls the underlying problem prompt rot. Both numbers are self-reported, but they are the vendor’s own caveats, which makes them more useful than the headline.

The gate sequence, and what each gate actually checks

The pipeline’s core is a sequence of gates where agents produce evidence and other agents, or humans, judge it.

Intake and diagnosis. Before any code changes, a Fellow agent must establish the case. From the post: “Before the Fellow fixes anything, it has to reproduce the problem, perform a differential diagnosis (to use hospital terminology), run whatever targeted investigation narrows it, and post a treatment plan: diagnosis, files to change, tests to add, risks, open questions. A Review Attending, a different agent instance running a prompt built around finding fault rather than fixing, reads the plan and approves, or rejects it. Only then does treatment start.”

The design choice worth noticing is the adversarial split: the reviewing agent is prompted to find fault, not to help. That is a prompt-engineering decision, and it is the kind that degrades quietly, which is presumably why prompt rot shows up as a named problem in the vendor’s own writeup.

Decomposition. For large work, a planning agent breaks it down first. In one example from the post, a single large GitHub issue became fifteen sub-issues with an explicit dependency graph, foundation first, before treatment began. This is where a chunk of those 1,299 self-filed issues comes from.

Escalation. This is the gate that most changes daily triage behavior. Per the post: “Every decision the human Chief makes that generalizes, is recorded in an append-only precedent log. An agent that wants to escalate has to read the log first, and if a precedent applies, it follows the precedent and cites it instead of escalating. Only a human can create precedents. Over five months, this has cut down how often agents escalate to the Chief.”

Note what is and is not claimed. The mechanism is precise: an append-only log, citation required, human-only writes. The outcome is a direction with no figure attached. “Cut down how often” could mean a 60% drop or a 5% drop, and the post does not say. The mechanism is borrowable; the reported effect is not yet a number you can plan around.

Discharge. The last gate before merge is an audit of process, not of code: “The last gate before merge does not re-review the code. It verifies that the review happened properly: an approval exists, the review template was filled out, no threads are unresolved, CI is green, the commit history is clean. It posts a checklist showing exactly what it verified, and if any box cannot be checked, discharge fails with that box left unchecked and annotated.”

This is a specific, reusable idea: separate “was the review done correctly” from “is the code correct,” and automate the former. It also makes explicit that the system trusts the earlier gates. If the Review Attending is compromised by a bad prompt, discharge will still pass, because discharge only checks that an approval exists.

What the independent evidence says about the fragile steps

Three pieces of independent research, none of them mentioned in the vendor post, bound what this model can safely delegate. All three are single-team studies and remain unreplicated, so treat them as bounds, not verdicts.

The intake assumption is the weakest link. The hospital model assumes a Fellow can reproduce a bug before diagnosing it. In the BugCraft evaluation of end-to-end crash bug reproduction for Minecraft, that step failed more often than it succeeded: “Even when provided with accurate step-by-step plans from the Step Synthesizer, the Action Model still failed to reproduce the bug in a majority of cases (56.14%).” The authors name the causes: poor decision making at 29.82% and agents stuck in loops at 19.30%. Minecraft crashes are not database bugs, and Cockroach Labs’ mature test infrastructure may make reproduction easier there, but the benchmark establishes that reproduce-before-treat is a gate agents fail at, not a formality.

Proactive discovery looks worse. If you want agents to find bugs without issue reports, the Active-SWE benchmark found “the best resolved rate reaching only 20.0%” for recorded bugs, with limited localization scores. The hospital pipeline starts from filed issues, so this does not contradict it directly. It does cap how far you can extend the model toward agents that patrol the codebase on their own.

The most useful result for a lead designing review process comes from a study of delegation contracts for AI coding agents. Formalizing what an agent must hand back changed reviewability without changing outcomes: “Objective outcomes saturated. All 64 runs passed the hidden acceptance checks; no run violated its authority boundary. On small, well-specified tasks, current models do not need a contract to succeed. Reviewability changed. Evidence sufficiency rose in 22 of 30 paired comparisons and fell in none (+0.83 on a 5-point scale; p < 0.0001”. The practical reading: on tasks agents already complete, structure does not make the fixes better, it makes the fixes checkable. That is an argument for adopting hospital-style gates as a review tool, and against crediting them with the million lines.

The economics, stated carefully

The same post reports that adding IBM Db2 support to the company’s MOLT tooling took under two days with no human-written code and a $4,172 token bill, described as 164x faster and 38x cheaper than an earlier Oracle support effort. Do not use that figure to justify the hospital pipeline. It is a different project, a different scope, a self-reported comparison against a 2024 effort, and exactly the genre of speed claim that circulates without scrutiny. It tells you something about the vendor’s enthusiasm for agent-written tooling and nothing verified about triage.

The cost you can actually reason about is the one the vendor discloses against its own system: roughly 29,000 tokens of instructions loaded per Discharge Nurse run, multiplied across fifteen agents and every case. Prompt-instruction overhead is a real line item in this architecture, and prompt rot is a maintenance liability: fifteen role prompts that must stay coherent with each other as the system evolves. A conventional triage process has meetings instead. Neither is free, and the post does not price one against the other.

Hospital model versus conventional triage: the operating checklist

The table below converts the five-month experiment into gates you can adopt independently of the vendor’s results. The “what to measure” column reflects the delegation-contract finding above: where output volume does not move, reviewability is the metric that can.

StageHospital model (as reported)Conventional triageWhat stays humanWhat to measure
IntakeFellow agent reproduces the bug and runs a differential diagnosisEngineer reproduces from the report, often skimmed under loadDeciding which reports enter the queueReproduction success rate; independent benchmarks put agent reproduction failure as high as 56.14% in hard environments
PlanTreatment plan posted before code: diagnosis, files, tests, risks, open questionsDesign note or PR description, quality varies widelyAccepting risk on plans that touch critical pathsEvidence sufficiency of plans; in one study, structure raised it in 22 of 30 paired comparisons
Plan reviewAdversarial Review Attending agent approves or rejectsSenior review, sometimes rubber-stamped under deadlineRulings where the agent reviewer and proposer disagreeRejection rate over time; a falling rate may mean better plans or a rotting prompt
EscalationAgent must cite an applicable precedent from the log or escalate to the human ChiefInterrupt-driven; the same senior answers the same question repeatedlyCreating precedents; only humans write to the logEscalation frequency; the vendor reports it fell over five months, with no figure published
DischargeGate verifies review happened: approval exists, template filled, CI green, history cleanMerge when CI passes and someone clicked approveSign-off on anything the checklist cannot verifyChecklist failure annotations as an audit trail

I would adopt the escalation row first, even with no agent involvement at all. An append-only precedent log that forces anyone, human or agent, to cite existing rulings before interrupting your most senior engineer pays for itself in a conventional process. The discharge row is second: auditing that review happened is cheap, and most teams discover their review process has holes only after an incident.

The apprenticeship gap

The vendor’s most candid paragraph is about the people the model fails. From the post: “But what about the new-grad who doesn’t have a full calendar, the software engineering judgment that decades of experience buys, and has a desire to deeply learn how to be a stronger engineer? How does the model work for them? In short, it doesn’t fully meet their needs. We’re currently working on enhancements to the hospital which will allow humans to take triaged issues and craft their own plans and have them reviewed by the hospital agents.”

The mechanism of the loss is worth spelling out. Bug triage and small fixes are historically where junior engineers build the diagnostic judgment the model now requires of reviewers. If agents absorb that work, your pipeline produces seniors who learned by reviewing agent plans, a skill that exists only if the plans are worth learning from, and juniors with no reps. The proposed remedy, humans writing plans that hospital agents review, is unshipped as of the post, so it is a stated intention, not a validated fix.

If you adopt the gates, decide explicitly which repair work stays reserved for juniors and treat that reservation as a cost of the staffing pipeline, not an inefficiency. The delegation-contract study offers one supporting hint: structured plans raised evidence sufficiency for reviewers, which is precisely the artifact a junior learns to read and write. Plan-writing practice is the part of the apprenticeship this model can preserve, if you route it to humans deliberately.

What to measure, and why this is one company’s word so far

The verdict holds with its conditions attached. Borrow the gates: reproduction and differential diagnosis before code, an adversarially prompted plan review, a human-only precedent log that damps escalations, and a discharge check that audits the review rather than re-doing it. Measure escalation frequency, plan rejection rates, and evidence sufficiency of the artifacts agents hand you, because those are the quantities the independent evidence says structure can move. Do not plan around the million lines, the 1,299 issues, or the reported drop in escalations, all of which are self-reported volume or direction claims from a single company with no published baseline, running on a mature codebase with strong tests and CI that most teams do not have.

Two open questions decide whether this model ages well. The first is whether a second organization replicates it on a less polished codebase; the Fellow gate is the most likely failure point, since reproduction is where independent benchmarks show agents failing most. The second is whether the apprenticeship patch ships and works, because a triage system that consumes junior-level work without producing senior judgment is borrowing against a training pipeline it does not replace. Until either is answered, the hospital is a well-documented experiment worth mining for parts, not a proven way to run maintenance.

Frequently Asked Questions

What specific gates should an engineering lead adopt from the hospital model?

Borrow the gates: reproduction and differential diagnosis before code, an adversarially prompted plan review, a human-only precedent log that damps escalations, and a discharge check that audits the review rather than re-doing it.

What metrics should be tracked instead of raw fix counts?

Measure escalation frequency, plan rejection rates, and evidence sufficiency of the artifacts agents hand you, because those are the quantities the independent evidence says structure can move.

How does the model affect junior engineer development?

Bug triage and small fixes are historically where junior engineers build the diagnostic judgment the model now requires of reviewers. If agents absorb that work, your pipeline produces seniors who learned by reviewing agent plans, a skill that exists only if the plans are worth learning from, and juniors with no reps.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. Cockroach Labs' account of the experimentcockroachlabs.comAccessed
  2. Active-SWE benchmarkarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy