groundy
Models & Research

mdlARC's 44% ARC-AGI-1 Claim: What the Small Budget Leaves Out

An author-reported 44% on ARC-AGI-1 public eval for about $0.67 of rented compute shows what small budgets can justify, and what still needs replication.

Published Updated 5 references
A compact copper furnace makes a mosaic key for one abstract lock while enormous general engines stand idle behind it.
On this page6 sections

A single author reports training a 75-million-parameter transformer from scratch to 44% on the ARC-AGI-1 public evaluation set, for about $0.67 of rented GPU time. The figures come from Mithil Vakde’s writeup and the mdlARC repository, not from an independent evaluation, and they are measured on the benchmark’s public split, where answers have circulated online for years. What the result can support is narrow and still interesting: on this benchmark, for systems that train on the test tasks themselves, a relatively small model and a cheap run may be worth testing before assuming that a large general-purpose model is necessary. What it cannot support is a headline about frontier models. The author’s own comparison rules exclude that matchup; he measures his system against others that train the way his does, and he argues LLM scores on this public split should not be counted at all.

The result, in the author’s numbers

The repository states the claims plainly: 44% on the ARC-AGI-1 public evaluation, a total compute cost of about $0.67, roughly two hours on a 5090 rented from vast.ai, and a standard transformer of 75M parameters. It calls the score one of the best non-LLM results available today and the cheapest by far at that performance. The blog post describes the training itself as 1.5 hours on the same GPU. Those are different quantities and worth keeping apart: 1.5 hours of training, about two hours of end-to-end compute for a script that runs training and inference together.

This is the second public iteration of the system. The earlier version reached 27.5% on the same public split for $1.8, in under three hours on an A100 rented through Google Colab; the author’s post headline rounded that to $2 and two hours. The previous writeup framed the run as 333 times cheaper to train than TRM and HRM, test-time-trained systems he notes require multiple H100s for multiple days. The new post reports that the model scores the same as TRM and HRM, beats many LLMs, and reaches 7% on ARC-AGI-2. Every number in this paragraph is the author describing his own system.

Public eval or private eval?

ARC-AGI-1 is 400 training tasks and 400 evaluation tasks, grid-in, grid-out puzzles designed as a test of fluid intelligence: a solver sees a few demonstration pairs per task and must produce exact output grids, with three trials allowed per test input. The dataset’s own guidance warns developers not to leak information from the evaluation set into their algorithm.

That warning is the load-bearing context for the 44%. The evaluation tasks are public and have been for years, which is why the author himself argues that LLM scores on the public leaderboard are worthless, because the answers are on the internet and frontier labs train on the internet. For LLMs, he writes, only private scores should count. The caution cuts in his direction too, and he takes a version of it: the repository includes an optional step that deletes the raw data, the solutions file, and the dataset-building scripts before the run, so a reproducer can demonstrate the system is not reading answers directly.

The materials available do not show a private-split score for the 44% model. The earlier post noted that the 27.5% version had not been tested on the private eval. Until a private or otherwise held-out number exists, 44% is a public-split result reported by the person who built the system.

What the 67 cents covers

$0.67 is the price the author paid for roughly two hours on a 5090 rented from vast.ai. It is not a reproducible total cost for the work. Rental rates move, the run assumes CUDA above 12.8, ideally above 13.0, plus torch, numpy, numba, matplotlib and flash-attn, and the figure excludes the dataset build, the author’s development time, the earlier $1.8 run, and any experiments that failed. What a reader can reproduce from the repository is the attempt: clone it, build the dataset, optionally delete the leakage-prone files, and run one of three modes (low, medium, high). Whether the attempt lands near $0.67 depends on the rental market that day.

The author also has a considered objection to how cost is compared on this benchmark. The ARC Prize leaderboard charts cost-per-task against performance, and he argues that axis misleads twice: it counts only online compute, ignoring the large offline pretraining behind LLM entries, and dividing by task count makes little sense for systems like his, TRM and HRM that train on all test tasks at once, so training cost amortizes across the whole set. His stated comparison set is models that do similar test-time training. That framing is defensible, and it is one more reason the result should not be translated into a claim about general-purpose models.

A specialist, not an assistant

mdlARC does one thing. The author describes the earlier version as a vanilla four-layer autoregressive transformer, no recursion, no handcrafted architecture, with one admitted exception: inference-time augmentation (AAIVR) and per-task embeddings, the part he calls not “bitter lesson” pilled and wants to remove. The system adapts to evaluation tasks at test time. Using a task’s demonstration pairs for adaptation is different from training on its hidden target answer: a reproduction must establish that separation. The author compares this approach with other test-time-trained systems, including TRM and HRM. Asked why it works, his answer is that he does not know yet; compression is his best guess, and ablations are pending.

That specialization matters for the comparison the headline invites and should not get. When the earlier post claimed the model beats every single non-thinking LLM in existence, with an edit clarifying he meant frontier labs’ released models rather than ARC finetunes, the author’s own later argument undercuts it: if LLM public-eval scores are contaminated, beating them on the public eval says little about relative capability. The newer post makes a comparison with other purpose-built test-time-trained systems. That is a more relevant comparison class, but its score, compute accounting and protocol still need checking.

What a small budget can justify

If the result replicates, it establishes something specific: in the test-time-training regime on ARC-AGI-1, strong public-split performance does not require large models or large budgets. A 75M-parameter transformer and about a dollar of rented compute are, on the author’s account, enough to match systems trained on multiple H100s for days. It is a claim about the compute needed for this particular benchmark procedure. The low stated rental cost makes an independent attempt attractive; it does not make that attempt optional.

Three practices from this episode travel beyond ARC:

  • Treat public-split scores on old benchmarks as weak evidence, for everyone. Published answers create a contamination risk. Deleting solution files can guard against direct reads during a run, but cannot by itself rule out answers influencing earlier training or development choices.
  • State the denominator on every cost claim. “Trained for $0.67” means two hours at one rental price for one run of one script. It does not mean the research cost $0.67, and it does not mean your reproduction will.
  • Compare within class. The official leaderboard separates reasoning systems, base LLMs and competition-grade entries, flags preview results as unofficial, and shows only systems that cost under $10,000 to run. Read those categories and the evaluation protocol before treating two plotted scores as comparable; the chart is a starting point for that check.

What still needs replication

Three things would move this from interesting to settled: an independent reproduction of the 44% on the public split at current rental prices, a private-split score under the benchmark’s verification policy, and the ablations the author says are coming. His own standard for LLMs, private scores only, is the right standard for his system too.

The useful next step is a reproduction with a recorded code revision, evaluation protocol, GPU rate and full run cost. A matching score would strengthen the claim that this specialized procedure can work cheaply. A private evaluation would answer a different question about generalization beyond a familiar public set. Keep those questions separate: neither an inexpensive run nor an impressive public score settles both.

References

Follow the links in the article for context. The supporting material is collected here for further reading.

  1. 44% on ARC-AGI-1 in 67 centsmvakde.github.ioAccessed
  2. mvakde/mdlARCgithub.comAccessed
  3. New Pareto Frontier on ARC-AGImvakde.github.ioAccessed
  4. fchollet/ARC-AGI: The Abstraction and Reasoning Corpusgithub.comAccessed
  5. ARC Prize Leaderboardarcprize.orgAccessed

Join the discussion

Share a useful perspective or ask a question about this article.

Discussion guidelinesComments privacy