arXiv:cs.LG· Zhishen Sun, Hongzhan Wang, Sizhe Dang, Guang Dai, Haishan Ye·· 4 小时前AI 评分35
DART-ES:用进化策略微调 LLM 的难度感知重加权与定向回放
DART-ES: Difficulty-Aware Reweighting and Targeted Replay for Fine-Tuning LLMs with Evolution Strategies
AI 导读
研究者提出 DART-ES,一种难度感知重加权与定向回放的进化策略(ES)微调方法,通过扰动种群的通过率估计各问题局部可解性,构建动态难度状态以同时指导难度重加权和稀有可解样本回放,无需额外难度模型或反向传播。
正文
Abstract:Evolution Strategies (ES) enable memory efficient full parameter fine-tuning of large language models (LLMs) using only forward computation. However, standard ES uniformly averages rewards across problems and compresses problem level population feedback into a single scalar, making it difficult to capture how the learning value of each problem changes with model capability. To address this limitation, we propose Difficulty-Aware Reweighting and Targeted Replay for Evolution Strategies (DART-ES). DART-ES estimates the local solvability of each problem from its pass rate across the perturbation population and aggregates historical observations to construct a dynamic difficulty state. This shared state jointly guides continuous difficulty reweighting and rare solvable sample replay, thereby improving perturbation direction evaluation and training data allocation without introducing an additional difficulty model or backpropagation. Extensive experiments show that DART-ES achieves good fine-tuning performance. DART-ES outperforms ES on all five base models and improves the average accuracy from 72.07\% to 73.53\%, exceeding the 73.26\% achieved by GRPO on GSM8K. Across five challenging mathematical reasoning benchmarks, DART-ES achieves an average accuracy of 49.20\%, compared with 48.34\% for ES and remains competitive with strong 7B models trained with RL. Further experiments show consistent gains in instruction tuning, code generation and the Countdown task with a 14B model, demonstrating strong generalization across tasks and scalability to larger models. Beyond performance gains, DART-ES also shows clear advantages in system efficiency. It reduces runtime per step by 15.2\%--50.2\% and peak memory usage per GPU by 21.1\%--51.1\% compared with GRPO. Despite performing full parameter fine-tuning, DART-ES also requires less runtime and GPU memory than GRPO+LoRA.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.06993 [cs.LG] |
| (or arXiv:2610.06993v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.06993 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Zhishen Sun [view email]
[v1]
Sun, 4 Oct 2026 07:42:11 UTC (278 KB)
来源:arXiv:cs.LG · arxiv.org