arXiv:2609.30057v1 Announce Type: new Abstract: LLM agents automate GPU kernel optimization by repeatedly composing candidates and measuring their duration on real GPUs. Existing systems preserve measurement fidelity by reserving a GPU for an entire agent session or benchmarking command. However, this results in poor utilization because only a small fraction of command execution requires exclusiv

Topological visualization of KREX: Concurrent Kernel Benchmarking on Shared GPUs via Region-Granular Exclusivity
Brave API

KREX (arXiv:2609.30057v1) is a runtime system that enables concurrent kernel benchmarking on shared GPUs through region-granular exclusivity, addressing the low utilization and measurement fidelity issues of existing session- or command-granular reservation methods. By allowing LLM agents to mark only the timing-sensitive critical regions of a benchmarking command as exclusive, KREX permits high off-region concurrency while ensuring accurate kernel duration measurements.

The system enforces exclusivity within these marked regions by draining in-flight GPU work, blocking new submissions, and freezing sibling host processes to prevent CPU contention, while reusing persistent GPU contexts outside these regions to minimize serialization overhead. Evaluated on NVIDIA and AMD GPUs, KREX delivers up to 3.4× higher benchmarking throughput compared to command-granular baselines, with negligible timing inflation (0.30% for kernels >10 ms, 1.58% for >1 ms, and 3.90% for >0.1 ms).

Key technical mechanisms include: In-region exclusivity: Isolates GPU execution and host CPU cores for the measuring threads to protect measurement fidelity against interference. Off-region concurrency: Maximizes utilization by allowing multiple agents to execute non-timing parts of their commands simultaneously. * Persistent Contexts: Avoids the high cost of repeated, node-wide serialized context creation by reusing GPU contexts in persistent processes.

This approach preserves the relative performance signals that guide agent search, with candidate ranking-flip rates remaining within three percentage points of uncontended repeats, making it suitable for large-scale automated kernel optimization workflows.

Generated 8d ago
Open-Weights Reasoning

KREX addresses a practical inefficiency in LLM-agent-driven GPU kernel optimization. Such agents iteratively generate candidate kernels and measure their runtime on real GPUs, but existing benchmarking systems often preserve timing fidelity by reserving the entire GPU for one agent session or one benchmarking command. The paper observes that only a small fraction of that command execution actually requires exclusive GPU access—typically the timing-critical kernel execution—while the remaining phases, such as compilation, setup, data movement, and post-processing, can often proceed concurrently with work from other sessions.

The key contribution is a concurrent benchmarking system that enforces exclusivity at a finer granularity: rather than treating a whole command as an atomic unit, KREX partitions benchmarking work into regions and protects only the regions where interference would corrupt timing measurements. This allows multiple kernel-optimization sessions to share a GPU while still preserving measurement fidelity for the critical kernels. The result is a shift from coarse GPU ownership to region-level scheduling and interference control, enabling higher utilization without sacrificing the reliability required for kernel performance comparison.

This matters because automated kernel optimization can involve many short, repeated benchmark iterations, and idle gaps from coarse-grained exclusivity can substantially reduce throughput and increase cost. By improving shared-GPU utilization, KREX can accelerate agent-based kernel search, make multi-tenant GPU use more practical, and provide a design pattern for systems that need both high measurement fidelity and high concurrency on expensive accelerator resources.

Generated 8d ago
Sources