跳到正文
arXiv:cs.LG· Gaurav Gupta, Vatshank Chaturvedi, Sudipta Sengupta, Jun Huan, Anoop Deoras·· 5 小时前AI 评分53

arXiv 论文提出 intent-execution gap,用 agent 轨迹剖析模型行为

Dissecting model behavior through agent trajectories

AI 导读

arXiv 论文(2606.17454)提出 intent-execution gap,即模型意图与 agent harness 实际执行之间的错配,并认为缩小该差距与工具和执行循环设计同等重要。

正文

View PDF HTML (experimental)

Abstract:AI agent performance is not just a modeling problem, it is fundamentally a systems problem. The advanced capabilities of models are realized through agent harnesses. Therefore, a gap between model assumptions and harness behavior can easily prevent the model's full capabilities from translating into agent performance. We formalize this as the `intent-execution' gap: the mismatch between what the model intends and what the harness executes, and vice versa. We argue that minimizing this intent-execution gap is as important as other aspects of harness design such as tools and execution loops. To illustrate the impact of this harness-model alignment, we develop a simple and customizable harness called `Simple Strands Agent' (SSA). SSA aims to find the bulk of common patterns which generalize across different model families (such as Claude, Gemini, GPT, Grok, Qwen), as well as a small number of model-specific preferences. We make two contributions: (i) we \textbf{reproduce or improve on the pass@1} performance reported by diverse model-provider families on popular agentic benchmarks (SWE-Pro, SWE-Verified, Terminal-Bench-2, and Deep-SWE), and (ii) building on an \textbf{analysis of 154\texttt{k} trajectories generated by SSA}, we look beyond the \texttt{pass@1} numbers which tend to be relatively even across frontier models. By representing agent trajectories in code state-spaces, we compute solution distance and backtracking to observe model-level differences in problem-solving behavior. Finer-grained metrics such as edit frequency, testing activity, and phase-transitions reveal how individual models allocate effort across different stages of problem solving.
Comments: 124 pages, 59 Figures, 24 Tables
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2606.17454 [cs.AI]
  (or arXiv:2606.17454v3 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2606.17454

arXiv-issued DOI via DataCite

Submission history

From: Gaurav Gupta [view email]
[v1] Tue, 16 Jun 2026 03:17:03 UTC (1,889 KB)
[v2] Wed, 17 Jun 2026 04:51:06 UTC (1,889 KB)
[v3] Thu, 1 Oct 2026 21:05:11 UTC (3,096 KB)

来源:arXiv:cs.LG · arxiv.org