arXiv:cs.LG· Kaushal Mhapsekar, Bita Aslrousta, Brijesh Kumar Bhayana, Paula Contreras, Azam Ghanbari, Ethan Goodman, Anna Andriiko, Samira Mirbagher Ajorpaz·· 4 小时前AI 评分36
CACHEFORGE:LLM 引导的端到端生成式缓存替换策略,兼顾性能与硬件效率
CACHEFORGE: LLM-Guided End-to-End Generative Cache Replacement Policy for Performance and Hardware Efficiency
AI 导读
CACHEFORGE 是首个将大语言模型嵌入硬件感知闭环、端到端演化缓存替换策略的框架:每轮由 LLM 生成新的 C++ 替换逻辑,在基于 trace 的 CRC-2 ChampSim 模拟器上评估,并通过奖励塑形、结构检查、动态变异、温度调度与跨策略交叉保证可行性。
正文
Abstract:Modern cache replacement designs saturate because they operate within fixed representational structures, hand-crafted and heuristic based feature-engineered predictors, or offline imitation models that cannot generate new decision logic on their own. At the same time, replacement is shaped by the causal interaction of prefetching, thrashing, spatial locality, and access-type behavior, producing an enormous design space that is difficult to traverse manually. Prior approaches typically rely on heuristics, parameter tuning, or imitation of an offline optimal policy, capturing correlations rather than synthesizing new mechanisms. As a result, their performance gains often plateau and they overfit under dynamic workload conditions.
CACHEFORGE is the first framework to evolve cache-replacement policies end-to-end by embedding a large language model inside a governed hardware-aware loop. In each iteration, the LLM proposes new C++ replacement logic, the policy is evaluated under a trace-based CRC-2 ChampSim simulator, and the framework enforces feasibility through reward shaping, structural checks, dynamic mutation, temperature scheduling, and cross-policy crossover. This closed-loop generation-evolution loop specifically designed for cache replacement policy enables the discovery of compact policies that satisfy hardware constraints while exploring algorithmic transformations beyond fixed predictor structures.
Across SPEC CPU2006, CACHEFORGE outperforms all CRC-2 baselines. It improves the total hit rate by 27.36%, 19.69%, 13.72%, 13.15%, 11.83%, and 5.73% over MPPPB, ReD, Hawk-eye, SHiP++, LIME, and LRU, respectively. On memory-intensive workloads, it increases IPC by 10.15%, 7.89%, 6.34%, 3.64%, 3.12%, and 2.71% over LRU, MPPPB, LIME, ReD, SHiP++, and Hawkeye.
| Subjects: | Hardware Architecture (cs.AR); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07668 [cs.AR] |
| (or arXiv:2610.07668v1 [cs.AR] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07668 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kaushal Mhapsekar [view email]
[v1]
Tue, 6 Oct 2026 03:03:13 UTC (1,511 KB)
来源:arXiv:cs.LG · arxiv.org