groundy
Industry & Business

AI Research Agents vs Kaggle Experts: What the Version Logs Show

TraceML pairs human and agent Kaggle logs to show agents collapse into narrow loops, supporting trajectory audits over leaderboard scores for autonomy decisions.

Published 7 references
A scuffed translucent green dinosaur holds a yellow resin egg and looks backward while standing on a circular trail of overlapping footprints. Hard shadows fall across the warm ivory background.
On this page12 sections

Teams deciding whether an auto-research agent can own a multi-hour ML project now have an unusual window into how agents actually work, not just what they score. TraceML, a new study pairing 430 human Kaggle trajectories with 207 agent runs on the same seven competitions, finds that both agent scaffolds tested collapse into narrow loops that final scores conceal. The practical consequence: gate autonomy on trajectory-level audits, not leaderboards.

What TraceML actually paired

Most agent evaluations report an outcome: a score, a medal, a success rate. TraceML instead reconstructs how the work unfolded. The corpus behind TraceML contains 4,465 human Kaggle trajectories across 134 competitions, expressed as version logs: each notebook or code version a competitor submitted, in sequence, annotated by the type of work it represents.

The comparison that matters sits in a paired subset of seven competitions worked by both humans and agents. On the human side, 430 trajectories spanning all author tiers. On the agent side, TraceML records 207 trajectories: 11 baseline runs of the Codex scaffold, 7 Codex runs carrying a planning prompt, and 13 MLEvolve searches linearized into root-to-leaf branches. Every agent run operated under a twelve-hour budget. The two scaffolds were chosen because their search topologies differ: Codex works as a sequential coding agent, while MLEvolve maintains an evolving tree of candidate solutions.

That pairing is what makes the study useful to a platform lead rather than only to benchmark authors. The humans and agents faced the same tasks under comparable version-level measurement, so differences in work pattern are not confounded with differences in the task set. The tradeoff is scale: seven competitions, two scaffolds, a single study. Keep that pinned to every conclusion below.

How experts spend a competition

The human profile that emerges from the version logs is one of alternation. According to TraceML, experts cycle through four kinds of work: data work (cleaning, features, augmentation), validation (building and stress-testing the evaluation split), model changes, and ensembling. Just as telling, they return to approaches they had set aside: top humans return on 9% of eligible versions per TraceML, and 78% of those returns end above the version they went back to.

This is the operational signature of long-horizon competence. Kaggle competitions reward exactly the behavior that makes any long ML project succeed: trying a direction, parking it without forgetting it, and recombining it later when the evidence shifts. The version log records not just what worked, but the management of a portfolio of half-tested ideas. If you are deciding whether an agent can own a project measured in days rather than prompts, this portfolio behavior is what you are buying, and it is invisible in any single final score.

Two scaffolds, two collapses

Against that baseline, both agent scaffolds degrade, and they degrade in opposite directions.

Codex spends its steps re-weighting ensembles and tuning submissions, per the TraceML results. It finds a working configuration early and then polishes it: adjust blend weights, nudge hyperparameters, resubmit. Direction never changes; the tuning never stops. The paper describes this as tuning without changing direction.

MLEvolve shows the mirror image. It mutates its model in place, churning through architectural and training changes without the consolidation phase that would lock in gains through validation work and ensembling. Changing direction without consolidating.

The shared failures matter more than the divergence. Neither scaffold pivots at the human rate, and neither reopens abandoned work. The shelved-idea return, the behavior that lets an expert’s version log function as a portfolio rather than a diary, is absent in both: across all agent runs, Codex returns once and MLEvolve never. Two different search topologies, one continuous and one tree-structured, landed on the same deficit from opposite sides. That is suggestive that the problem lies in how these agents manage work over hours, not in one scaffold’s implementation detail. It is still only two data points.

Why the leaderboard misses it

If both scaffolds collapse, why do agent benchmark results keep looking respectable? Because outcome-only scoring cannot see the difference between a healthy search and a degenerate one.

Independent evidence for this blindness comes from GCPC, a checklist-based scoring framework for agent trajectories, studied on SkillsBench v1.1 with 87 tool-using tasks across eight domains and 20,818 labeled trajectories from 234 harness-model-condition configurations. GCPC measures which task requirements are satisfied and how much progress an agent makes toward completion, and its authors show that this captures execution differences invisible under task-level success rates, including differences between runs with identical final outcomes. Two trajectories can end at the same score while one searched and the other looped. The leaderboard records them as equal.

A coding-agent study makes the same point from the feedback side. The Observability Gap shows that in a Blender-based 3D scene generation task, output-only human feedback fails to produce full-scene success even when the agent can rediscover the core utility functions on its own. Supervising an agent by looking only at what it finally produces is not a weaker version of supervising its process; it is a different activity that misses different failures.

Together these results explain the trap in outcome-based autonomy decisions. A team grants an agent more independence because its scores look fine, the scores look fine because they aggregate away the collapse, and the collapse is exactly what a twelve-hour unattended run amplifies. Aggregate slopes are where process problems hide, a pattern Groundy has seen before when decomposing usage into individual trajectories overturned a population-level conclusion.

The trajectory-audit checklist

TraceML’s schema converts directly into an audit any team can run on its own agent logs. Before granting an agent unattended hours, require a sample of its runs on your tasks and score three behaviors against a human baseline from your own engineers’ version history:

Audit metricHuman profile (per TraceML)Collapse signature to flag
Direction-change rateRegular pivots between data, validation, model, and ensemble workNear-zero pivots (tuning loop) or churn without consolidation
Validation share of stepsSustained fraction of versions devoted to validation workValidation crowded out by submission tuning or model mutation
Shelved-idea reuseAbandoned approaches reopened and recombined laterNo revisit of any shelved line of work

Two cautions apply to using this table. First, the metrics are derived from one study’s schema; they expose the failure modes TraceML observed, but nobody has demonstrated that passing all three predicts downstream project success. Treat them as tripwires, not certifications. Second, any automated version of this audit is judge-sensitive: GCPC’s authors swapped the scoring judge to Claude Haiku 4.5 and watched agreement drop from Spearman ρ=0.916 on outcome-inclusive scores to ρ=0.615 on trajectory-level scores. Trajectory grading by a model needs calibration against human review before you trust its numbers.

When a planning prompt helps, and what it cannot fix

The strongest counter-evidence to a pessimistic reading is inside TraceML itself. The authors distilled human practice into a short planning prompt and gave it to Codex. The result: the behaviors the prompt named moved toward the human profile, and scores improved. The collapse is partly instruction-addressable, not a fixed property of the model.

The limit is in the same sentence of the paper: the effort profile stayed agent-shaped. Instruction closed only the part of the gap that reduces to instructions. Naming a behavior in a prompt got the agent to perform more of it; it did not reproduce the underlying portfolio management that makes the behavior useful. Notably, the prompt moved only what it named, which means each additional failure mode a team discovers requires its own explicit patch, and the ones nobody has named yet remain unpatched.

Routing advice follows. For bounded tasks where the failure modes are known and enumerable, a planning prompt distilled from your own experts’ practice is cheap and worth doing. For genuinely open-ended, multi-day research ownership, a prompt is a partial mitigation, and the residual risk sits precisely in the unobserved, unnamed behaviors that trajectory review exists to catch.

Who grades the grader

Everything above assumes a trusted reviewer reads the trajectories. That assumption needs its own audit.

CrossAudit frames the problem bluntly: an AI scientist should not grade its own homework, yet in the systems its authors examined, the reviewing agent usually comes from the same model family as the producing agent, or at least the same vendor. Model evaluators are known to favor their own generations. When the graded benchmark doubles as the agent’s feedback signal, the bias compounds: the agent optimizes toward what its sibling prefers.

The obvious fix, cross-vendor review, reduces but does not eliminate the problem. CrossAudit notes that frontier models plausibly train on largely overlapping public corpora, so two different vendors’ models can be confidently wrong together wherever the shared literature is wrong. Independence of the grader from the producer is necessary and insufficient.

There is also a reproducibility gap underneath the benchmarks teams might cite to justify trusting an agent. A pilot audit of twelve LLM agent benchmark papers found only two cleared the full-disclosure threshold for inference settings, where a missing run date counts against papers that use closed APIs, and zero of the eight agent benchmark papers scored pinned their environments with content-addressed digests. Repository tags can be re-pushed against different content; digest pinning is what makes a claimed environment checkable. A benchmark result you cannot re-run is an anecdote with a table.

The operational cost of trajectory review

Shifting trust from scores to trajectories is not free, and the honest accounting belongs in the autonomy decision.

A September 2026 perspective on agent swarms as researchers estimates that weeks of agent production generated months of human review work. If your review process is the gate, its throughput caps how much agent output you can responsibly accept, and an agent that produces faster than you can audit is not autonomous; it is unaudited. Sampling strategies help, but they reintroduce exactly the aggregate-level blindness the trajectory audit was meant to remove, just at a coarser grain.

Long unattended runs also import supply-chain exposure that output review will never see. A large-scale study of agent skill ecosystems found that 26.1% of 31,132 analyzed skills contained at least one vulnerability across 14 patterns, a survey of trustworthy agentic AI reports: data exfiltration at 13.3%, privilege escalation at 11.8%, and prompt injection. Those figures reach this article secondhand, through the survey’s citation of the underlying study (Liu et al. 2026, arXiv:2601.10338); the primary paper has not been checked directly here. A trajectory audit that covers only modeling decisions misses what the agent installed and invoked along the way. For runs measured in hours with network access, the audit needs a dependency and tool-invocation column too.

Verdict and limits

The defensible position for an ML platform lead today: do not grant unattended long-horizon ownership on the basis of final scores. Gate it on a trajectory audit (direction-change rate, validation share, shelved-idea reuse) measured against your own engineers’ version logs, keep the grader outside the producing model’s vendor, treat planning prompts as partial mitigations, and size the agent’s output budget to your review capacity.

Now the asterisks, which are load-bearing. TraceML is a single author-reported study with no independent replication, covering two scaffolds on seven Kaggle competitions under a twelve-hour budget. It demonstrates behavioral collapse under those conditions; it does not establish a ceiling on what current or future scaffolds can do, and Kaggle competition work is not all ML work. The three audit metrics have no demonstrated predictive link to project success, only to exposing failure modes that scores miss. Automated trajectory graders are judge-sensitive (ρ=0.916 falling to 0.615 on a judge swap), cross-vendor review still shares correlated errors from overlapping training corpora, and the benchmark literature teams might lean on is frequently not reproducible at the environment level. The version logs justify distrust of leaderboard-only autonomy decisions. They do not yet justify confidence in any particular audit as a certification of trust.

Frequently Asked Questions

How does a planning prompt affect agent performance in TraceML?

The authors distilled human practice into a short planning prompt and gave it to Codex. The result: the behaviors the prompt named moved toward the human profile, and scores improved. The collapse is partly instruction-addressable, not a fixed property of the model.

What is the main limitation of using automated trajectory graders?

Second, any automated version of this audit is judge-sensitive: GCPC’s authors swapped the scoring judge to Claude Haiku 4.5 and watched agreement drop from Spearman ρ=0.916 on outcome-inclusive scores to ρ=0.615 on trajectory-level scores. Trajectory grading by a model needs calibration against human review before you trust its numbers.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. TraceMLarxiv.orgAccessed
  2. The Observability Gaparxiv.orgAccessed
  3. CrossAuditarxiv.orgAccessed
  4. A pilot audit of twelve LLM agent benchmark papersarxiv.orgAccessed
  5. a survey of trustworthy agentic AI reportsarxiv.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy