跳到正文
arXiv:cs.AI· Yunbo Long, Guangya Hao, Yuhan Liu, Yiting Duan, Longyan Tan, Yunchen Long, Hao Wu·· 6 小时前AI 评分40

编码智能体的自我纠错该携带多少证据?自适应 Dirichlet 证据用于自我蒸馏

How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation

AI 导读

研究者提出 Effective-Evidence Self-Distillation(EESD),将执行反馈中的相对转移支持与有效伪计数质量分开表示,再用 Dirichlet 后验生成带不确定性惩罚的权重用于 KL 锚定的纠错学习,在四组模型-领域历史扫描中把可见观测从 1 增至 8,使未来结果 NLL 下降 55.0-59.3%。

正文

View PDF HTML (experimental)

Abstract:Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.
Subjects: Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
Cite as: arXiv:2610.08514 [cs.AI]
  (or arXiv:2610.08514v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08514

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yunbo Long [view email]
[v1] Tue, 6 Oct 2026 15:17:56 UTC (3,291 KB)

来源:arXiv:cs.AI · arxiv.org