跳到正文
arXiv:cs.LG· Daifeng Li, Huiqiang Jiang, Chengruidong Zhang, Wei Wu, Xudong Guo, Jianhong Tu, Jianwei Zhang, Binhang Yuan, Dayiheng Liu·· 3 小时前AI 评分45

D2K-Bench:LLM 智能体能否把专家设计转化为高效 GPU Kernel?

D2K-Bench: Can LLM Agents Turn Expert Designs into Efficient GPU Kernels?

AI 导读

研究者推出 D2K-Bench,一个包含 26 项任务、85 个工作负载的诊断性基准,用于衡量 LLM 智能体将专家设计指导转化为高效 GPU Kernel 的能力。

正文

View PDF HTML (experimental)

Abstract:GPU kernels generated by large language model (LLM) agents can remain less efficient than expert implementations, but runtime alone does not reveal how the gap relates to design discovery and implementation. We introduce D2K-Bench, a diagnostic benchmark of 26 tasks and 85 workloads that measures how effectively agents translate expert design guidance into efficient GPU kernels. The guidance covers L1: high-level algorithmic insights, L2: dataflow design, and L3: low-level optimization tricks, including dependencies among these levels. Pairwise runs with and without guidance share task descriptions, workloads, tools, hardware, and a 350-turn budget. Complementary assessments examine independently proposed designs and the design properties implemented in generated code. Across five models on NVIDIA B200 GPUs, guidance raises correctness over 130 model-task pairs from 93.1% to 98.5% and increases the Performance Score over all 26 tasks from 1.46 to 1.95. For the three frontier models with correct submissions on all 26 tasks in both runs (GPT-6-Astra, Claude-Opus-4.8, and GPT-5.6-Sol), geometric mean speedup increases from $1.69\times$ to $2.49\times$. Across all five models, the mean combined implementation score increases from 57 to 70 out of 100. These results show the value of expert design guidance while identifying design properties that remain unimplemented.
Comments: 30 pages, 4 figures
Subjects: Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as: arXiv:2610.03226 [cs.LG]
  (or arXiv:2610.03226v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.03226

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Daifeng Li [view email]
[v1] Fri, 2 Oct 2026 12:42:12 UTC (252 KB)

来源:arXiv:cs.LG · arxiv.org