跳到正文
arXiv:cs.AI· Yaopei Zeng, Congchao Wang, JianHang Chen, Nan Wang, Yurui Chang, Lu Lin·· 6 小时前AI 评分42

Critic Experience Bank:为 LLM 智能体实现自演进步骤级置信度估计

Critic Experience Bank: Self-Evolving Step-Level Confidence Estimation for LLM Agents

AI 导读

研究者提出 Critic Experience Bank(CEB),一个无需训练即可让 LLM 智能体做步骤级置信度估计的框架,将已完成轨迹的回溯反馈转为可复用证据。CEB 在四个智能体基准、三个 critic 主干共十二种设置下取得最佳或并列最佳的 ECE、Brier score 和 AUC,ECE 最高较最强免训练基线降低 53.8%。其置信度分数还能提升选择性执行与模拟任务成功率。

正文

View PDF HTML (experimental)

Abstract:LLM agents operate in stateful environments, where a single erroneous step can waste limited interaction budget or cause irreversible effects before task failure becomes apparent. Reliable deployment therefore requires step-level confidence estimation: estimating, before execution, the probability that a proposed action will advance the task. Existing LLM confidence estimators are typically designed for static question answering under a fixed task context and evaluation criterion. For an agent, however, its action productivity depends on an environment transition that is observed only after execution. To address this challenge, we introduce Critic Experience Bank (CEB), a training-free framework that turns feedback from completed trajectories into reusable evidence for future confidence judgments. After each trajectory, an LLM assigns hindsight productivity pseudo-labels to individual actions and stores them with the critic's original pre-execution confidence, task context, action, and observed feedback. For a new action, by retrieving related productive and unproductive experiences, CEB grounds pre-execution confidence in feedback from completed trajectories to condition a fixed LLM critic. CEB thereby adapts over a task stream without parameter updates or ground-truth step labels at deployment. Across four agent benchmarks spanning offline and live web navigation, mobile GUI and shell tasks, and three critic backbones, CEB achieves the best or tied-best ECE, Brier score, and AUC in all twelve benchmark-backbone settings under rule-based step labels, reducing ECE by up to 53.8% relative to the strongest training-free baseline. Its confidence scores also improve downstream utility in selective execution and simulated task success.
Comments: 20 pages, 5 figures
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2607.12397 [cs.AI]
  (or arXiv:2607.12397v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2607.12397

arXiv-issued DOI via DataCite

Submission history

From: Yaopei Zeng [view email]
[v1] Tue, 14 Jul 2026 06:17:09 UTC (346 KB)
[v2] Tue, 6 Oct 2026 00:26:43 UTC (349 KB)

来源:arXiv:cs.AI · arxiv.org