arXiv:cs.AI· Linh-An Phan, MingXue Wang, Guangyu Wu, Feng Pan, Zhaoyu Pang, Yanbin Zhang·· 5 小时前AI 评分42
LiteTrajEval:面向生产级 AI 智能体的轻量级评分标准引导轨迹评估
Lightweight, Rubric-Guided Trajectory Evaluation for Production AI Agents
AI 导读
LiteTrajEval 是一种预算受限的轻量级轨迹评估架构,离线提取领域规则画像,在线预处理轨迹、标记启发式失败信号并在固定预算下序列化,再由单个评分标准引导的 LLM 裁判生成结构化诊断报告。
正文
Abstract:Trajectory evaluation is essential for improving the reliability of LLM-based agents, but production use makes it expensive to run repeatedly. Modern agents generate long traces containing tool calls, observations, retries, and external outputs, while not all raw tokens are equally useful for diagnosis. We present \textit{LiteTrajEval}, a lightweight architecture for budget-bounded trajectory evaluation. LiteTrajEval derives compact domain-specific rule profiles offline, then preprocesses each trajectory online, marks heuristic failure signals, serializes it under a fixed global budget, and invokes a single rubric-guided LLM judge to produce structured diagnostic reports. Evaluated on public Magentic-One-style and $\tau$-bench-style trajectory datasets, LiteTrajEval improves failure-localization alignment with human annotations by roughly 20--35 percentage points on Magentic-One and up to 23 percentage points on $\tau$-retail compared with AgentRx, while reducing cost by about 6$\times$ and evaluation time by more than 8$\times$. This solution has also been deployed in our enterprise agentic platform.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.03315 [cs.AI] |
| (or arXiv:2610.03315v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03315 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mingxue Wang [view email]
[v1]
Fri, 2 Oct 2026 13:51:06 UTC (352 KB)
来源:arXiv:cs.AI · arxiv.org