跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Stefan Abi-Karam, Miaoyan Zhou, Callie Hao·· 14 小时前AI 评分36

SyntheticHLS:用 LLM 构建多样化合成高层次综合数据集

SyntheticHLS: Building Diverse Synthetic High-Level Synthesis Datasets using LLMs

AI 导读

SyntheticHLS 是一个用 LLM 生成大规模、复杂且多样化合成 HLS 数据集的框架,其迭代反馈引导变异循环可将种子设计逐步转化为更复杂、可扩展的设计。

正文

View PDF HTML (experimental)

Abstract:Deep learning and large language models (LLMs) are rapidly gaining adoption in semiconductor design, driving demand for training datasets. Most efforts focus on hardware description languages (HDLs) while designs for high-level synthesis (HLS), a popular approach to domain-specific accelerators, remain scarce. HLS dataset efforts emphasize manual curation or design parameterization, seldom addressing high-quality LLM-based generation or diversity in code length, hierarchy, design-space size, latency, resource utilization, and application domain, potentially limiting model generalization.
We propose SyntheticHLS, a framework for generating large-scale, complex, diverse synthetic HLS datasets using LLMs. Its two key ideas are: 1) an iterative feedback-guided mutation loop that uses paired HLS source code and design-space specifications to incrementally transform seed designs into more complex, scalable designs; and 2) quantitative metrics of HLS design complexity and design-space scalability that serve as measurable objectives for LLM-guided mutation.
We systematically cross-validate an HLS Quality-of-Results (QoR) deep learning model trained and tested across common HLS benchmarks, zero-shot synthetic designs, and iteratively mutated synthetic designs. Synthetic designs transfer well to common benchmark test sets while the reverse does not hold. SyntheticHLS's iteratively mutated designs provide the most generalizable training corpus among the datasets studied. Analysis of the mutation process and dataset shows that metric-guided trajectories consistently improve targeted complexity and scalability objectives without regressing non-target metrics. Mutated designs span a substantially broader, more diverse design space than zero-shot generated designs.
Our framework, dataset, and evaluation are open-source: this https URL.
Comments: Accepted and to be presented at the International Conference on Field Programmable Technology (FPT) 2026
Subjects: Hardware Architecture (cs.AR); Machine Learning (cs.LG)
Cite as: arXiv:2610.00106 [cs.AR]
  (or arXiv:2610.00106v1 [cs.AR] for this version)
  https://doi.org/10.48550/arXiv.2610.00106

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Stefan Abi-Karam [view email]
[v1] Wed, 9 Sep 2026 02:40:29 UTC (3,330 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org