arXiv:cs.AI· Yunbo Long, Guangya Hao, Yuhan Liu, Yiting Duan, Longyan Tan, Yunchen Long, Hao Wu·· 6 小时前AI 评分40
编码智能体的自我纠错该携带多少证据?自适应 Dirichlet 证据用于自我蒸馏
How Much Evidence Should a Coding Agent's Self-Correction Carry? Adaptive Dirichlet Evidence for Self-Distillation
AI 导读
研究者提出 Effective-Evidence Self-Distillation(EESD),将执行反馈中的相对转移支持与有效伪计数质量分开表示,再用 Dirichlet 后验生成带不确定性惩罚的权重用于 KL 锚定的纠错学习,在四组模型-领域历史扫描中把可见观测从 1 增至 8,使未来结果 NLL 下降 55.0-59.3%。
正文
Abstract:Execution feedback lets coding agents revise programs and learn from their own corrections. A correction's learning weight should reflect both the transitions supported by its executions and the amount of evidence behind that support. We introduce Effective-Evidence Self-Distillation (EESD), which represents these quantities separately. Normalized execution relevance determines relative transition support and an effective pseudo-count mass; a Dirichlet posterior then produces an uncertainty-penalized weight for KL-anchored correction learning. Under a symmetric prior, changing mass preserves category ordering, and effective mass yields a supervised coefficient bounded by its matched fixed-mass counterpart. Across four model-domain history sweeps, increasing visible observations from one to eight reduces future-outcome NLL by 55.0-59.3%. At eight observations, effective mass achieves lower NLL than fixed mass in all four comparisons. In the primary matched DeepSeek/RunBugRun study, argmax predictions agree on all 3,000 examples, with the largest NLL gain under concentrated relevance. After one correction-learning round, DeepSeek/CodeARC all-tests Pass@1 increases from 15.0% to 20.4%, with a paired 95% source-bootstrap interval of [+2.8, +8.0] percentage points. The twelve-setting downstream evaluation establishes the model-domain scope of this update. These results show how separating evidence support from evidence mass changes probability estimation and correction learning in coding agents.
| Subjects: | Artificial Intelligence (cs.AI); Software Engineering (cs.SE) |
| Cite as: | arXiv:2610.08514 [cs.AI] |
| (or arXiv:2610.08514v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08514 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yunbo Long [view email]
[v1]
Tue, 6 Oct 2026 15:17:56 UTC (3,291 KB)
来源:arXiv:cs.AI · arxiv.org