arXiv:cs.AI(全量分类)· Shuyao Xiao, Shengling Wang, Xuan Chen, Ke Chao, Ming Cui, Feifei Qian, Chaoyang Mei, Fanlin Meng, Ziming Yu, Junxi Yin·· 5 小时前AI 评分36
LLM 智能体恢复机制何时失效:CIR 因果评估方法
When Harnesses Lose the Signal: Causal Evaluation of Recovery in LLM Agents
AI 导读
研究将 LLM 智能体的恢复机制建模为因果决策问题,提出轻量策略 Causal Intervention Router(CIR),在恢复前判断干预是否值得。在 Qwen3-14B 执行长程 ALFWorld 任务时,CIR 将成功率从 70.33% 提升至 73.33%,增益 3.00 个百分点,且未改动任何观测正确的轨迹。对照实验表明,恢复带来的收益无法仅由环境返回的新观测解释。
正文
Abstract:Large language model agents rely on external harnesses to pass information between the model and its environment and to recover from execution errors. Yet recovery is usually judged only by average task success. This hides an important tension. The same operation can rescue a failing trajectory or disrupt one that would otherwise succeed. We frame recovery as a causal decision problem. Starting from the same execution state, we compare what happens with and without recovery, separate rescue from harm, and study how the value of recovery changes over time. We then introduce the Causal Intervention Router (CIR), a lightweight policy that uses information available before recovery to decide when intervention is worthwhile. On long-horizon ALFWorld tasks with Qwen3-14B, CIR raises success from 70.33% to 73.33%, a gain of 3.00 percentage points. It leaves all evaluated trajectories with correct observations untouched. Additional controls show that the benefit of recovery cannot be explained solely by the new observation returned by the environment. These results provide a practical way to evaluate recovery and apply it selectively.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.00372 [cs.AI] |
| (or arXiv:2610.00372v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00372 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shuyao Xiao [view email]
[v1]
Wed, 30 Sep 2026 07:04:25 UTC (463 KB)
来源:arXiv:cs.AI(全量分类) · arxiv.org