arXiv:2609.30059v1 Announce Type: new Abstract: Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kern

Topological visualization of KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
Brave API

KernelOPT is a multi-agent system designed to optimize GPU kernels by respecting the structural decisions of compilers like PyTorch Inductor, rather than treating models as black boxes. It utilizes five profiling-guided LLM agents to selectively optimize only the generated Triton sub-kernels, while preserving vendor library calls such as cuBLAS and cuDNN.

The system employs a rigorous four-gate verification cascade—covering static validation, multi-seed correctness, model-level float64-fallback verification, and performance gating—to ensure that any optimized kernel is both correct and faster than the compiler baseline. If no candidate passes all gates, the system safely reverts to the compiler baseline, ensuring zero regressions.

Evaluated on 250 KernelBench problems, KernelOPT achieves significant geometric mean speedups over torch.compile: 1.40× speedup on Level 1 problems (51/100 optimized). 1.15× speedup on Level 2 problems (31/100 optimized). * 1.07× speedup on Level 3 problems (12/50 optimized).

Key optimizations include operator fusion (e.g., fusing five operations into one kernel) and targeted improvements on pointwise and reduction operations, demonstrating that LLM-driven optimization is most effective when it leverages compiler-generated structure rather than overriding it.

Generated 8d ago
Open-Weights Reasoning

Problem and approach. KernelOPT targets a central bottleneck in modern deep learning systems: compiler-generated GPU kernels, such as those produced by PyTorch Inductor, often leave substantial performance headroom compared with expert-tuned kernels. The paper’s focus is not merely on making individual kernels faster, but on optimizing kernels in the context of how they are dispatched and executed within a real model. This distinction matters because end-to-end performance is affected by launch overhead, fusion boundaries, memory access patterns, synchronization, and interactions among neighboring operators—factors that are easily missed when optimizing kernels in isolation. The proposed method uses an LLM-driven agentic search loop to generate, modify, and evaluate kernel candidates, with the search conditioned on dispatch-level execution behavior rather than only on standalone kernel correctness or raw latency.

Key insight and significance. The main contribution is to reframe GPU kernel optimization as a dispatch-aware, workload-level search problem. Instead of treating a kernel as a self-contained code snippet, KernelOPT evaluates candidate kernels according to how they perform when integrated into the broader execution path of a model. This allows the optimizer to favor changes that improve practical runtime behavior—such as reducing dispatch costs, improving memory efficiency, or better aligning kernel shapes with hardware constraints—rather than chasing local speedups that may not translate into end-to-end gains. The work is significant because it addresses a common failure mode of LLM-assisted kernel optimization: strong results on isolated kernels but limited gains in full inference or training pipelines. By coupling agentic code generation with dispatch-aware feedback, the approach offers a more realistic path toward automated GPU kernel optimization that can complement compiler tooling and reduce reliance on manual expert tuning.

Generated 8d ago
Sources