arXiv:cs.LG· Ajay Vohra, Tao Chen, Neeti Narayan, Caron Zhang·· 3 小时前AI 评分41
DeReAct:为可靠 AI 智能体分解推理与行动
DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents
AI 导读
DeReAct 是一种模块化智能体架构,将动作授权与完成控制外化为两个门控策略:Critic 在执行前验证提议动作,Context Manager 重建环境支持的 State 并认证任务完成。
正文
Abstract:ReAct-based agents typically rely on a single LLM policy to propose actions, interact with the environment, and decide when a task is complete. This coupling makes action authorization and completion control difficult to enforce independently, allowing errors to propagate and unsupported completion claims to terminate execution. We introduce DeReAct, a modular agent architecture that externalizes two gating policies: a Critic that validates proposed actions before execution, and a Context Manager that reconstructs an environment-supported \textsc{State} and certifies task completion.
Across GAIA and SWE-bench Verified, DeReAct improves Pass@1 most for weaker Brain models, with gains of 6.5--7.0 points for Qwen3-Coder-480B and 4.2--5.2 points for Claude Sonnet~4.5; gains diminish as Brain capability increases. Trajectory and ablation analyses show that external gating is effective when targeted failures are sufficiently prevalent and the gating policy is itself sufficient. With Claude Opus~4.5, Pass@1 remains comparable to ReAct, while DeReAct produces more evidence-complete and constraint-satisfying trajectories, indicating that completion control can trade earlier termination for stronger grounding. Overall, DeReAct improves weaker agents while retaining grounding benefits as models strengthen.
| Subjects: | Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.02351 [cs.AI] |
| (or arXiv:2610.02351v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02351 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tao Chen [view email]
[v1]
Thu, 1 Oct 2026 18:25:07 UTC (407 KB)
来源:arXiv:cs.LG · arxiv.org