跳到正文
arXiv:cs.CL· Hengrui Gu, Xiaotian Han, Kaixiong Zhou·· 3 小时前AI 评分41

WRIT:面向多轮用户交互智能体的写读密集型轨迹合成方法

WRIT: Write-Read Intensive Trajectory Synthesis for Multi-Turn User-Facing Agents

AI 导读

WRIT 是一种沿写决策数量与单次决策证据负担两个复杂度轴合成多轮智能体训练轨迹的方法,先构造写密集与读密集任务,再多样化用户行为指令并在可执行环境中模拟交互生成完整轨迹。

正文

View PDF HTML (experimental)

Abstract:Multi-turn user-facing agents must infer user intent from incomplete requests, collect missing information through dialogue and tools, and execute valid actions. A training trajectory records this process as an interleaved sequence of user messages, agent responses, tool calls, etc. Synthesizing sufficiently complex trajectory has become a central route to train agents: existing pipelines often increase difficulty by composing multiple user requests into longer tasks, producing write-intensive trajectories that train sequential execution.
We argue that a single write decision can itself be difficult when the agent must gather and compare substantial read-tool evidence before its arguments become identifiable, a challenge that write-intensive data alone cannot address. Guided by this insight, we propose WRIT (\uline{W}rite-\uline{R}ead \uline{I}ntensive \uline{T}rajectory Synthesis), a pipeline for synthesizing multi-turn agent training trajectories along two complexity axes: the number of write decisions in a task and the evidence burden of each individual decision. WRIT first generates write-intensive and read-heavy tasks. It then diversifies user behavior instructions to reflect realistic conversational variation, and finally simulates agent-user interactions in an executable environment to produce complete training trajectories. The resulting data trains agents not only for longer task execution, but also for robust, evidence-grounded decision making under high information load. With only 2K synthesized trajectories, a 4B model trained on WRIT outperforms GPT-5.1 no-think on $\tau^2$-bench and substantially reduces inference-time token usage, showing that compact SFT data can convert part of expensive test-time reasoning into efficient agent behavior.
Comments: EMNLP 2026 Main Conference
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2606.02908 [cs.CL]
  (or arXiv:2606.02908v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2606.02908

arXiv-issued DOI via DataCite

Submission history

From: Hengrui Gu [view email]
[v1] Mon, 1 Jun 2026 21:25:06 UTC (4,023 KB)
[v2] Wed, 7 Oct 2026 15:09:37 UTC (4,033 KB)

来源:arXiv:cs.CL · arxiv.org