跳到正文
arXiv:cs.AI· Toby D. Pilditch, Konstantinos Voudouris, Alexandra Abbas, Cozmin Ududec·· 6 小时前AI 评分43

Transect:为长视野 LLM 智能体评测保留可观测性

Transect: Retaining Observability for Long-Horizon LLM Agent Evaluations

AI 导读

开源工具包 Transect 基于 Inspect Scout 构建,可将长视野智能体运行中的事件、token 用量、子智能体活动和模型生成的行为标签对齐到同一轮次时间线,并支持导出底层数据表。

正文

View PDF HTML (experimental)

Abstract:Frontier AI evaluations increasingly use open-ended, agentic, long-horizon tasks whose transcripts can span hundreds of pages of outputs and actions from complex multi-agent networks. The observability envelop-the range of what evaluators can reliably infer about an agent's behaviours-is therefore narrowing. Language model assistants can help classify and interpret agent behaviour but also afford human evaluators significant analytical degrees of freedom, threatening the reproducibility and auditability of language-model-based transcript analysis. Transect is an open source package built on Inspect Scout to help evaluators understand how a long agent run unfolded, identify behaviour worth investigating, and check interpretations against the transcript. Users specify task context and behavioural vocabulary in a reusable evaluation-family configuration, with judge models and analysis settings supplied separately. Transect's navigable reports align recorded events, token use, sub-agent activity, and model-generated behavioural labels on a common turn-based timeline. Reviewers can quickly grasp a run's narrative, trace any label or event to its source turns, and export the underlying data tables for cross-run analysis. We demonstrate the workflow on an AI R&D evaluation that generated almost 13 million tokens, dividing the agents' work into behavioural phases aligned with research-skill classifications, sub-agent delegations and interactions, and token use. The combined view shows a focus on operational work and manuscript production, with little evidence of a sustained hypothesis generation stage-arguably a necessary component for high-quality scientific outputs. Transect's flexible, customisable transcript-analysis pipeline will enable evaluators to keep pace with longer, more complex, more frequent AI evaluations while supporting scientific rigour, transparency, and reproducibility.
Comments: 27 pages, 5 figures
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08364 [cs.AI]
  (or arXiv:2610.08364v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08364

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Toby Pilditch [view email]
[v1] Tue, 6 Oct 2026 13:52:15 UTC (216 KB)

来源:arXiv:cs.AI · arxiv.org