跳到正文
arXiv:cs.AI· Kavienan Jegatheesan, Gayathri Lihinikaduarachchi·· 6 小时前AI 评分49

工具会撒谎:工具反馈被污染时数学智能体的可靠性

When Tools Lie: Reliability of Mathematical Agents Under Corrupted Tool Feedback

AI 导读

研究通过受控污染框架,在 31 道数学题上测试隐藏拦截器把工具调用结果替换为貌似合理的错误信息时智能体的表现。无验证时准确率从 100% 骤降至 72.4%,强制同上下文反思可完全恢复至 100%,可选验证仅在模型主动调用时才有提升。显式检测后重启问题在 100% 情况下成功恢复。

正文

View PDF HTML (experimental)

Abstract:Mathematical problem solving often requires deterministic computational steps that agents delegate to tools and implicitly trust. Yet tools can fail silently, returning plausible but incorrect results. How well can agents detect and correct corrupted tool call outputs? We study this through a controlled corruption framework where a hidden interceptor replaces tool call results with plausible incorrect information on targeted problems. We evaluate agents across 31 problems under four verification designs including no verification (baseline), mandatory same-context reflection, optional fresh-context verification, and optional structural verification. Without verification, corruption causes dramatic accuracy loss, from 100% down to 72.4%. Mandatory reflection fully recovers this performance to 100%. Optional verification improves accuracy only when models actively invoke it. Our results show that checking frequency is strongly associated with robustness differences, while unequal invocation prevents a controlled comparison of verifier quality. A supporting recovery experiment shows that full problem restart succeeds in 100% of cases after explicit detection. These findings demonstrate that verifier availability and verification policy are separate components of mathematical-agent reliability. Mandatory policies enforce verification while optional policies depend on the model's own choice to invoke it.
Comments: 8 pages, 2 figures, 3 tables. Accepted to the 6th Workshop on Mathematical Reasoning and AI (MathAI) at NeurIPS 2026
Subjects: Cryptography and Security (cs.CR); Artificial Intelligence (cs.AI); Software Engineering (cs.SE)
Cite as: arXiv:2610.08097 [cs.CR]
  (or arXiv:2610.08097v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.08097

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Kavienan Jegatheesan [view email]
[v1] Tue, 6 Oct 2026 10:28:19 UTC (1,115 KB)

来源:arXiv:cs.AI · arxiv.org