跳到正文
arXiv:cs.LG· Akshay Manglik, Vijay S. Kalmath, Jason Qin, Apaar Shanker, Kaustubh Deshpande, Yash Maurya, Veronica Chatrath, Levi Lentz, Yuan Xue·· 6 小时前AI 评分46

Insights Generator:面向 LLM 智能体的语料级执行轨迹诊断系统

Insights Generator: Systematic Corpus-Level Trace Diagnostics for LLM Agents

AI 导读

针对 LLM 智能体失败诊断仍依赖人工抽查少量执行轨迹的问题,研究者提出 Insights Generator(IG),一个通过跨轨迹语料提出并验证假设来生成有证据支撑的诊断报告的多智能体系统。

正文

View PDF HTML (experimental)

Abstract:Diagnosing failures in LLM agents remains largely manual. Practitioners inspect a small subset of execution traces, form ad-hoc hypotheses, and iterate. This process misses patterns that only emerge across trace populations and does not scale to production corpora where individual traces span tens of thousands of tokens. We formalize the problem of corpus-level trace diagnostics. Given a corpus of execution traces, the goal is to produce grounded natural-language insights that characterize systematic behavioral patterns across trace groups, each linked to supporting evidence. We present the Insights Generator (IG), a multi-agent system that answers diagnostic questions by proposing and testing hypotheses across the trace corpus to produce an evidence-backed insights report. We evaluate IG across qualitative and objective dimensions, spanning rubric-based report assessment and downstream performance improvements achieved by implementing IG insights. Human experts using IG reports improve scaffold performance by 30.4 pp over the unmodified baseline scaffold, and coding agents leveraging IG-derived insights show consistent and stable gains. Across benchmarks, IG's scout-investigator architecture produces findings comparable in detection coverage to competing approaches, while domain experts rated IG reports as leading depth and evidence quality.
Comments: v4: expanded evaluation with the APEX-Agents benchmark and a GPT-5.4 cross-family judge. Revised analysis and appendices
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG); Software Engineering (cs.SE)
Cite as: arXiv:2605.21347 [cs.AI]
  (or arXiv:2605.21347v4 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2605.21347

arXiv-issued DOI via DataCite

Submission history

From: Veronica Chatrath [view email]
[v1] Wed, 20 May 2026 16:13:53 UTC (1,429 KB)
[v2] Thu, 21 May 2026 16:51:51 UTC (1,429 KB)
[v3] Thu, 4 Jun 2026 23:27:42 UTC (1,429 KB)
[v4] Tue, 6 Oct 2026 20:07:31 UTC (973 KB)

来源:arXiv:cs.LG · arxiv.org