跳到正文
arXiv:cs.AI· Wenlong Zhang, Zhengbo Jiao, Chenxu Zhang, Lekang Jiang, SiYuan Ma, Qituan Zhang, Guo Chen, Linfeng Zhang·· 5 小时前AI 评分43

递归 Harness 自我改进:面向前沿推理数据合成

Recursive Harness Self-Improvement for Frontier Reasoning Data Synthesis

AI 导读

研究者提出"任务-工具协同演化"框架,实现推理数据合成中的递归 Harness 自我改进(RSI),在生成过程中将求解失败转化为可复用技能,并在每批任务后修订技能、提示词与工作流。在数学、编程和科学领域,经过十四轮演化,求解器平均准确率从 100.0% 降至 54.8%。用 10K 合成数学样本微调的 27B 学生模型在 APEX 上达到 62.5% mean-16 准确率。

正文

View PDF HTML (experimental)

Abstract:Generating progressively harder reasoning problems requires synthesis procedures that adapt as the task distribution evolves. Existing task-level recursion reuses generated problems as seeds but leaves the construction harness unchanged. We present task-harness co-evolution, a framework for recursive harness self-improvement (RSI) in reasoning-data synthesis. Online self-improvement converts intermediate solver failures into reusable skills during generation. Post-task self-improvement revises skills, prompts, and workflows after each batch, adopting candidates only when they generate harder valid tasks within a bounded cost increase. Model weights and verification criteria remain fixed. Across mathematics, coding, and science, mean solver accuracy decreases from 100.0% to 54.8% over fourteen evolution rounds. Ablations show that combining both update schedules produces harder tasks than fixed-harness recursion or either schedule alone. The resulting data improves downstream SFT and GRPO performance. In particular, a 27B student fine-tuned on 10K synthesized mathematics examples achieves 62.5% mean-16 accuracy on APEX, competitive with selected frontier-model references. These results support adapting the synthesis harness alongside the tasks to generate increasingly challenging data with downstream training value.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.03548 [cs.AI]
  (or arXiv:2610.03548v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.03548

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Wenlong Zhang [view email]
[v1] Fri, 2 Oct 2026 16:30:15 UTC (1,525 KB)

来源:arXiv:cs.AI · arxiv.org