arXiv:cs.AI· Babak Barazandeh, Connor Swanson, Chinmay Kulkarni, Nikhil Mungel·· 3 小时前
OnTrack:基于流式结构感知最优传输的 LLM 智能体轨迹实时监控与干预
OnTrack: Real-Time Monitoring and Intervention in LLM Agent Trajectories via Streaming Structure-Aware Optimal Transport
AI 导读
研究者提出 OnTrack,一种流式监控机制,通过将智能体步骤及其依赖与历史成功运行对比,在约每步 1 毫秒内告警或阻断智能体。在 SWE-bench 轨迹上,基于前 8 步,该方法将失败轨迹排在成功轨迹之下的 AUROC 比内容相似度方法高 0.057,配合中止策略可节省约 18% 的失败运行算力,其中 83% 的中断运行确实走向失败。
正文
Abstract:Agents are deployed in applications from trip planners and stock trading to IT incident triage. In most cases, LLM agents work autonomously with minimal rule-based safeguarding, leading to cost and safety issues from irreversible actions. Recent works resolve this either by using a safeguard agent to monitor behavior or evaluating logs post-hoc. The first adds cost and latency to every step; the second delivers its verdict after the run, when tokens are burned and damage is done. To overcome this, we propose OnTrack, a streaming monitoring mechanism that compares an agent's steps and dependencies against recorded successful runs to alert users or block the agent in about a millisecond per step. We study this problem in three regimes of decreasing access: full reference access (historical runs and tool schemas), intermediate access (only tool schemas), and no prior knowledge (only step logs as generated). Expectation of OnTrack's monitoring capabilities reduces as data access drops, ranging from plan violation detection to identifying loops, stalls, and repeated tool calls. Finally, we evaluate OnTrack using SWE-bench trajectories. Based on the first 8 steps, our method ranks failing trajectories below succeeding ones better than content similarity approaches (+0.057 AUROC). With an abort policy, we save about 18% of compute that would be burned on failing runs, where 83% of interrupted runs were actually heading to failure (5 out of 6 aborts were correct).
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12375 [cs.AI] |
| (or arXiv:2610.12375v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12375 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Babak Barazandeh [view email]
[v1]
Thu, 8 Oct 2026 17:32:08 UTC (37 KB)
来源:arXiv:cs.AI · arxiv.org