Systems that generate GPU kernels with language models report high correctness rates. Those rates come from a single loose test: run the kernel on a few random inputs at one fixed shape and accept it if the output is close to a reference. A kernel can pass that test and still be silently wrong. It can return an ordinary number where the true answer is a NaN or an infinity, differ from run to run,

Topological visualization of A Contract-Grade Verifier for LLM-Generated GPU Kernels, and a Native Blackwell Backward for the Gated-Linear-Recurrence Family
Brave API

A Contract-Grade Verifier for LLM-Generated GPU Kernels is a research paper by Rishi Shah and Rishav Shrestha (arXiv:2608.12700, submitted August 13, 2026) that challenges the reliability of current benchmarks like KernelBench. The authors argue that standard correctness tests, which rely on running kernels on a few random inputs at a fixed shape, create a "correctness illusion" where broken kernels are falsely accepted.

To address this, the researchers developed a contract-grade verifier consisting of twelve adversarial gates that check for properties such as shape polymorphism, determinism, non-finite value propagation, and precision regimes. When applied to 2,638 machine-generated kernels previously deemed correct, the verifier found that 39.5% were broken in ways that standard tolerance thresholds could not excuse, and 62.1% carried at least one violation.

The paper also presents a second contribution: the first native Blackwell tcgen05 training backward for the gated-linear-recurrence (GDN) family. This kernel was independently verified against a double-precision oracle and passed the contract-grade verifier, serving as a positive control to demonstrate the tool's rigor. The study suggests that adopting tolerance-free contracts would significantly close the gap between reported and actual correctness in LLM-generated GPU code.

Generated Aug 31, 2026
Open-Weights Reasoning

The material targets a reliability gap in LLM-generated GPU kernels: reported correctness rates often rest on a single loose test—run a few random inputs at one fixed shape and accept the kernel if the output is close to a reference. That protocol can miss serious failures, such as a kernel returning a finite value where the reference is NaN or Inf, producing nondeterministic results across runs, or silently violating shape, dtype, or boundary conditions. The paper introduces a contract-grade verifier that treats generated kernels as small APIs with explicit preconditions, postconditions, and numerical contracts, and evaluates them against a broader suite of shapes, data types, edge-case values, and repeated executions.

Its key insight is that correctness for GPU kernels should be defined more rigorously than pointwise closeness on benign inputs. The verifier checks semantic properties that matter in real training and inference stacks: proper propagation of NaN/Inf, deterministic behavior, correct handling of boundary and degenerate cases, and consistency across repeated launches. This gives a much stronger signal about whether an LLM-generated kernel is safe to compile, deploy, and use in production, rather than merely plausible on a narrow benchmark.

The paper also contributes a native Blackwell backward kernel for the gated-linear-recurrence family, a class of sequence models whose forward pass is efficient but whose backward pass can be difficult to express efficiently with generic autograd code. By targeting NVIDIA Blackwell hardware with a fused, architecture-aware implementation, the work aims to reduce memory traffic and improve utilization for the backward pass, which is often the bottleneck in training these models. Together, the two contributions argue that future kernel-generation systems need both high-performance, hardware-specific kernels and trustworthy verification infrastructure—because a kernel that is fast but subtly wrong is worse than one that is slower but provably reliable.

Generated Aug 31, 2026
Sources