arXiv:cs.AI· Geetha Prasuna Yarramneni, Surya Selvam, Wilfried Haensch, Anand Raghunathan·· 7 小时前AI 评分35
TraceDSE:面向异构边缘 SoC 的 AI 推理工作负载智能体式设计空间探索
Agentic Design Space Exploration for Joint Hardware Configuration Selection and Mapping of AI Inference Workloads on Heterogeneous Edge SoCs
AI 导读
TraceDSE 是一种智能体式设计空间探索流程,可在异构边缘 SoC 上联合完成 AI 推理工作负载映射与 PU 配置选择。
正文
Abstract:Modern edge Systems-on-Chip (SoCs) integrate heterogeneous processing units (PUs) such as CPUs, GPUs, and NPUs, each with distinct performance and energy characteristics. Deploying AI inference workloads on them under real-time latency and energy constraints requires jointly mapping workloads to PUs and configuring each PU (e.g., selecting the number of active cores and the operating frequency). This joint space grows combinatorially, making exhaustive search infeasible. Most prior work on design space exploration (DSE) applies black-box optimization (BBO) such as evolutionary search, where each evaluation returns only aggregate metrics such as latency and energy. Recent LLM-guided DSE relies on the same sparse feedback. We observe that this limits its efficiency: it offers no insight into the design space or the reasons a design choice performs the way it does, and it leaves the reasoning abilities of LLMs largely unused. We present TraceDSE, an agentic DSE flow that performs joint workload mapping and PU configuration selection for AI inference on heterogeneous SoCs. TraceDSE is an iterative proposer-critic loop driven by richer feedback in the form of system execution traces. The LLM proposer agent generates candidate mappings and PU configurations for hardware evaluation. The LLM critic agent, equipped with programmatic trace-analysis tools, analyzes the traces to identify bottlenecks and suggest targeted refinements. This loop yields deeper insight into each design point, higher-quality decisions, and a more effective search. Across four AI inference workloads (models of varying complexity and a multi-model pipeline) on an Intel Meteor Lake SoC, TraceDSE consistently outperforms two state-of-the-art BBO tools, improving Pareto frontier hypervolume by up to 35% over NSGA-II and up to 68% over Bayesian optimization, while requiring ~6-9x fewer hardware evaluations.
| Comments: | Published at the 2026 ACM/IEEE International Symposium on Machine Learning for CAD (MLCAD '26), Jeju Island, Republic of Korea, September 7-9, 2026 |
| Subjects: | Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.07191 [cs.AR] |
| (or arXiv:2610.07191v1 [cs.AR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07191 arXiv-issued DOI via DataCite (pending registration) |
|
| Journal reference: | Proceedings of the 2026 ACM/IEEE International Symposium on Machine Learning for CAD (MLCAD '26), September 7-9, 2026, Jeju Island, Republic of Korea. ACM, New York, NY, USA, 8 pages |
| Related DOI: | https://doi.org/10.1145/3831599.3840344
DOI(s) linking to related resources |
Submission history
From: Geetha Prasuna Yarramneni [view email]
[v1]
Mon, 5 Oct 2026 18:08:58 UTC (1,794 KB)
来源:arXiv:cs.AI · arxiv.org