LLM-generated GPU kernels can occasionally clear a production bar today, but only after gates that isolated benchmarks do not run, and mainly for narrow, well-covered operations. That is the practical reading of FastKernels, a preprint benchmark surfaced in feeds on October 5, 2026, together with a cluster of sibling benchmarks that measure what happens when generated CUDA and Triton kernels face realistic acceptance criteria. Under those criteria, the numbers are sobering: on RealisticTritonBench’s new-kernel tasks, only about 20% of generated kernels pass all unit tests, and KernelBench-Verified reports that 28% of the best model’s correctly-generated kernels increase peak GPU memory relative to PyTorch.
One caveat belongs up front, because it shapes everything below. Every source cited here is a single-team preprint with no independent replication as of October 2026, and FastKernels’ numbers are author-reported. What the paper does provide is the most direct evidence yet on the production question, because its authors seeded the benchmark with five agents across 6,900 agent-hours and reported what survived: kernel-level speedups of up to 6.6× shrink to at most 1.25× end to end, and only 20% of winning kernel sets run correctly as-is (source). That supports a sharper question than “can LLMs write kernels”: what does a generated kernel have to clear before it belongs in a serving path, and which current results suggest any model clears it?
What FastKernels actually measures
Most kernel-generation benchmarks score a candidate kernel in isolation: compile it, run it against a reference on some synthetic shapes, compare latency. FastKernels takes a different position on what “production” means. According to the paper, “FastKernels also scores every candidate end to end, inside the models its tasks come from and on their production execution path (continuous batching, chunked prefill, CUDA graphs, torch.compile, tensor parallelism).” That is the design choice that matters for an inference engineer. A kernel that wins a microbenchmark can still lose inside a real serving stack, because CUDA graphs constrain memory allocation patterns, chunked prefill changes the shapes the kernel sees, and continuous batching means batch composition varies at runtime.
The task universe is also broader than prior efforts. The paper describes “a benchmark of 384 tasks drawn from 47 representative architectures across 8 categories, whose kernels suffice to reimplement 94.6% (472/499) of HuggingFace Transformers architectures with outputs matching the native implementations” (abstract). Read that number carefully: 94.6% is a statement about benchmark coverage, not a model score. It says the task set spans most of what HuggingFace Transformers ships. It says nothing about how well any LLM performs on those tasks; the agent results are two sections down.
The default scoring configuration is worth understanding because it encodes the paper’s own view of a production bar: “The default set is the 11-model subset on which agents are scored, because a single long-horizon agent already spends up to 2,500 agent-hours on it” (FastKernels). That set spans all 8 categories, TP 1–4, and BF16/FP16/FP8/MXFP4 across 48 workloads, and its models run 133 distinct L1, L3 tasks. Two details stand out. First, tensor parallelism from 1 to 4 and four numeric formats are treated as part of correctness, not extras. Second, the 2,500 agent-hour figure hints at the real cost of production-grade evaluation; running this kind of scoring is itself an infrastructure investment, which matters when you decide whether to build an in-house harness.
What winning kernels do inside a real serving stack
FastKernels seeded its leaderboard with five agents, Dr. Kernel, Claude Code, KDA, AKO and Codex, for a combined 6,900 agent-hours on the default set. Four of them reached kernel-level geomean speedups of 1.6–6.6× over the kernels production frameworks ship. End to end, the gains mostly evaporate: even after dropping every kernel that crashed or corrupted its model, deployable subsets span 0.96–1.25×, with LLM serving staying within 6% of the production kernels (source).
Swapped in unchanged, the picture is worse: “only 22 of the 110 winner-set deployments run correctly (13–58% of model families per agent)” (source). The drop analysis is the most useful part for anyone building a gate. “Compilation-stack conflicts (e.g. embedding kernels calling raw C++ extensions that torch.compile cannot trace, rotary kernels that specialize dynamic batch shapes) and crashes on engine shapes the kernel bench never produced account for 86% of the drops with an identified cause” (source). Every one of those failures is invisible to an isolated kernel check, which is the paper’s point and the reason its end-to-end tier exists.
Composition adds a third failure layer: “37% of the agents’ L3 winners break once their own lower-level kernels are active” (source). A layer-level kernel that passes on its own can fail when the primitive beneath it is also generated. Regenerating higher-level kernels on frozen lower-level winners recovers most of the loss, yielding 80 correct L3 winners across agents instead of 45.
And kernel-level scores mis-rank agents: “Claude Code matches or beats KDA at every level in isolation, yet KDA scores 3× higher end to end” (source). The gap is deployability, not speed: all 11 of KDA’s models stay valid versus 6 of Claude Code’s, and KDA gets there by wrapping the same Claude model in a planning and review harness that a kernel-level leaderboard cannot credit. If you are choosing between agents rather than building a gate, that end-to-end number is the one that matters.
Why earlier kernel benchmark scores don’t transfer
If your team has been watching kernel-generation results over the past year, you have seen optimistic numbers. The authors of robust-kbench explain why those numbers diverge from production outcomes: existing benchmarks carry exploitable loopholes and insufficient diversity, which is why they introduce “a new benchmark for rigorous evaluation of kernel performance and correctness across varied scenarios.” A benchmark whose test inputs are predictable, or whose correctness check runs on one shape at one dtype, invites solutions that pass the check rather than solve the problem.
The reward-hacking concern is concrete enough that CUDA-Harness builds its evaluation pipeline around it: “To dilute reward hacking in Text2CUDA, we construct Synthesis-Based Verification to provide isolated test data and progressive validation.” The mechanism matters more than the name. If the same pipeline (or the same model) that writes the kernel also produces the test data, the generator can learn the tests rather than the semantics. Decoupling test-data synthesis from kernel generation, then validating progressively, is the countermeasure. Any acceptance process your team builds should assume the generator has effectively seen your public test suite and hold out inputs accordingly.
How generated kernels fail today
The strongest counter-evidence to an unqualified “yes, LLMs can write production kernels” comes from benchmarks that score along multiple dimensions at once. On RealisticTritonBench, which evaluates Triton generation inside real AI frameworks, the best model tested (Qwen3.5-397B-A17B) reached only a 25.81% task success rate under stricter realistic criteria. The gap between passing tests and succeeding at the task is the benchmark’s central finding: across all task types the average unit-test pass rate is 60.33% and the full-test pass rate is 43.23%, yet overall task success is only 18.71%, because kernels that pass tests can still degrade model accuracy or add end-to-end latency in a deployed system. The failure pattern on novel work is sharper, as RealisticTritonBench reports:
Across the 11 New-kernel tasks, on average, only 20% of the generated kernels are able to pass all unit tests, which is the lowest among all categories. Moreover, even among the limited set of correct kernels, the end-to-end latency metrics are generally below 1.0, and the average numerical robustness (NR) is only 40%
Three failure modes are packed into that one finding (source), and each maps to a gate your harness needs. Correctness fails first (one in five kernels passing unit tests). Among the survivors, speed versus the reference fails (latency metrics below 1.0 means slower than baseline). And numerical robustness is its own failure mode: among those correct kernels the average NR is 40% on these new-kernel tasks, and RealisticTritonBench scores it as a model-level numerical-stability check, separate from unit tests. A gate that stops at unit tests or latency never measures this axis at all.
The fourth failure mode is memory, and it is invisible to latency-only scoring. KernelBench-Verified reports that “28% of correctly-generated kernels by the best model (GPT-5.5) increase peak GPU memory usage relative to PyTorch.” Most generated kernels do not increase peak memory, but the savings erode with scope: “memory savings shrink sharply as problems grow more complex, from 82% of single-operator problems to 36% of full-model architectures” (source). Peak memory is a production constraint, not a benchmark nicety: it determines batch size headroom and whether a kernel fits alongside KV cache under continuous batching. A kernel that wins on latency but raises peak memory enough to force a smaller batch can be a net throughput loss, and a latency gate alone will never catch that.
What a real acceptance gate looks like
The most useful template in this research set comes from a contest, not a benchmark. The FlashInfer kernel contest describes an acceptance process where “Candidate code was not accepted by assertion: it had to pass harness-controlled correctness, compilation, profiling, and latency gates.” The key phrase is “harness-controlled.” The candidate does not get to report its own numbers. Compilation is a gate, because plenty of generated code fails to build under production compiler flags. Profiling is a gate, because wall-clock latency can hide pathological behavior that profiling exposes.
Test inputs deserve the same rigor as scoring logic. Liger-Kernel, an author-reported Triton kernel library for LLM training, describes its practice: “For input shapes in testing, we use actual dimensions/hyper-parameters from the training process, such as a batch size of 44, a hidden dimension of 20482048, and a variable sequence length” (testing methodology). (The doubled digit appears in the source text.) FastKernels levels the same critique at its predecessors, which it says “evaluate kernels in isolation, with synthetic inputs and weak baselines, rewarding sandbox speedups that break or vanish in real inference systems,” and it builds its own kernel-level cases from the shapes its capture pipeline records while the models actually run. If your harness tests shapes your serving stack never emits, it is measuring the wrong thing.
Pulling these sources together, the acceptance criteria that recur across them form a checklist. This synthesis is our inference from the combined evidence, not a result any single paper reports:
| Gate | What to check | Evidence behind it |
|---|---|---|
| Compilation | Builds under production flags and toolchain | FlashInfer harness compilation gate |
| Correctness | Unit tests across real shapes, dtypes, TP 1–4 | FastKernels default set; RealisticTritonBench ~20% unit-test pass rate on new-kernel tasks |
| Numerical robustness | Model-level numerical-stability check passes; kernel maintains model accuracy | RealisticTritonBench reports ~40% average NR on new-kernel tasks |
| End-to-end latency | Faster than reference inside the real model, under continuous batching, chunked prefill, CUDA graphs, torch.compile | FastKernels production-path scoring |
| Deployment survival | The winning kernel set runs correctly as-is inside the real model | FastKernels: 22 of 110 winner-set deployments run correctly unchanged |
| Peak memory | Parity or better versus the native implementation | KernelBench-Verified memory regressions |
| Harness integrity | Test data isolated from generation; held-out inputs | CUDA-Harness Synthesis-Based Verification; robust-kbench loopholes |
| Realistic inputs | Actual serving dimensions, not synthetic shapes | Liger-Kernel testing practice |
The bar, and which kernels can clear it
The decision this evidence supports: gate every LLM-generated kernel on production-path acceptance before it touches a serving path. That means unit tests plus numerical robustness across the shapes, dtypes, and tensor-parallel configurations you actually run; end-to-end latency measured inside the real model under your serving stack’s execution mode; and peak-memory comparison against the implementation you would otherwise ship. Anything less is measuring benchmark fitness, and the robust-kbench authors have documented how gameable benchmark fitness is.
Expect kernels to clear this bar mainly for narrow, well-covered op classes: fused elementwise operations, normalization variants, and similar patterns that appear repeatedly across the 47 architectures FastKernels draws from. The inference here is ours, but it follows directly from the measured pattern that failure rates rise with task novelty (RealisticTritonBench’s new-kernel category being the worst), with problem complexity (KernelBench-Verified’s memory savings collapsing from single operators to full-model architectures), and with composition (FastKernels finding that 37% of L3 winners break once an agent’s own lower-level kernels are active). Whole-model kernel rewrites remain a research activity, not a deployment path.
There is movement on the generation side worth tracking. DICE, a diffusion LLM for CUDA kernel generation, trains with “a hierarchical progression from easy to hard across both the CUDA kernel data and task complexity,” a curriculum design that targets exactly the novelty gap where current models fail. If that approach holds up, the set of tractable op classes will widen. But “holds up” is doing work in that sentence: it is one more single-team preprint.
If generated kernels do start clearing production gates for narrow op classes, the economics of inference optimization shift in a specific direction. Scarce GPU-engineer time moves from writing kernels to building and maintaining correctness harnesses, because the harness becomes the asset that determines whether generation is safe. That shift does not make generation cheap: FastKernels reports that a single long-horizon agent spends up to 2,500 agent-hours generating kernels on the default set, and that each additional long-horizon agent costs about 2,000 agent-hours and 8–22B tokens. The harness is the option that converts an unbounded review burden into a bounded one.
The honest limitation, one more time: nothing here is independently replicated, and all eight sources are author-reported preprints. What FastKernels adds to the checklist is a documented failure taxonomy rather than a model ranking: 86% of its identified end-to-end drops came from compilation-stack conflicts and unseen engine shapes (source), failures an isolated kernel check cannot see. The checklist does not depend on any model’s score, which is why it is the durable part of this story. Model rankings will change within months. The gates a kernel must clear before it serves traffic will not.
Frequently Asked Questions
What is the end-to-end speedup of LLM-generated kernels in FastKernels?
End to end, the gains mostly evaporate: even after dropping every kernel that crashed or corrupted its model, deployable subsets span 0.96–1.25×, with LLM serving staying within 6% of the production kernels (source).
How many LLM-generated kernels pass unit tests on new-kernel tasks?
Across the 11 New-kernel tasks, on average, only 20% of the generated kernels are able to pass all unit tests, which is the lowest among all categories.
What percentage of generated kernels increase peak GPU memory?
KernelBench-Verified reports that “28% of correctly-generated kernels by the best model (GPT-5.5) increase peak GPU memory usage relative to PyTorch.”

Join the discussion
Share a useful perspective or ask a question about this article.