arXiv:cs.LG· Aheli Poddar, Sanskar Prasad, Arindam Samanta, Subha Chakraborty, Vishal Goyal, Rohit Singh Rathaur·· 4 小时前AI 评分42
KernelOPT:面向 GPU 内核优化的调度感知智能体搜索
KernelOPT: Dispatch-Aware Agentic Search for GPU Kernel Optimization
AI 导读
KernelOPT 是一个多智能体系统,将编译后的模型视为结构化产物,保留 cuBLAS、cuDNN 等厂商库调用,仅针对生成的 Triton 子内核,由五个性能分析引导的 LLM 智能体优化。
正文
Abstract:Deep learning inference and training performance depends critically on GPU kernel efficiency. Modern compilers such as PyTorch Inductor automatically generate GPU kernels from high-level model code, but frequently underperform expert-written implementations by wide margins. Recent LLM-assisted kernel optimizers can close this gap for standalone kernels, yet treat compiled models as black boxes, generally optimizing individual standalone kernels without respecting the compiler's structural decisions or verifying the model end-to-end. We present KernelOPT, a multi-agent system that treats compiled models as structured artifacts. It preserves vendor library calls (cuBLAS, cuDNN) and exclusively targets generated Triton sub-kernels using five profiling-guided LLM agents. A four-gate verification cascade applies static validation, multi-seed correctness checking, model-level float64-fallback verification, and performance gating ($\gamma{=}1.03$) to filter candidates and verify the re-stitched model end-to-end. When candidates fail verification, the system preserves the compiler baseline. The system accepts PyTorch nn Modules, standalone Triton kernels, and Helion kernels. Evaluated on 250 KernelBench problems (100 Level 1, 100 Level 2 and 50 Level 3) on NVIDIA H200, KernelOPT achieves geometric mean speedups over torch compile of 1.40$\times$ (L1), 1.15$\times$ (L2), and 1.07$\times$ (L3) across all kernels, including fallback cases. Optimized-only geomeans (excluding cases where verification gates preserve the compiler baseline) are substantially higher: 2.54$\times$ (L1: 36/100), 1.84$\times$ (L2: 23/100), and 1.37$\times$ (L3: 11/50), reflecting where the optimizer achieves meaningful leverage.
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2609.30059 [cs.DC] |
| (or arXiv:2609.30059v2 [cs.DC] for this version) | |
| https://doi.org/10.48550/arXiv.2609.30059 arXiv-issued DOI via DataCite |
Submission history
From: Aheli Poddar [view email]
[v1]
Thu, 24 Sep 2026 16:17:52 UTC (7,781 KB)
[v2]
Tue, 6 Oct 2026 17:23:07 UTC (7,786 KB)
来源:arXiv:cs.LG · arxiv.org