跳到正文
arXiv:cs.AI· Justin Chih-Yao Chen, Elias Stengel-Eskin, Yan Chen, Pol Llado, Scott Counts, Mohit Bansal, Benjamin Van Durme, Harsh Jhamtani, Gaurav Verma·· 6 小时前AI 评分46

TeleTune:从离线遥测数据中演化智能体技能

TeleTune: Evolving Agent Skills From Offline Telemetry

AI 导读

TeleTune 是一款从离线日志中学习文本技能库的框架,无需记录目标、无需回放即可优化,并能处理任务交错的轨迹。在 WorkArena 和 Online-Mind2Web 上,TeleTune 平均成功率分别达 77.1% 和 80.6%,较最强基线提升 6.7% 和 7.7%;在 WorkArena 训练数据最重扰动下仍保持 68.5%,高于最强基线 6.3%。

正文

View PDF HTML (experimental)

Abstract:Computer-use agents need to capture procedural knowledge of how people use software. User telemetry offers a scalable source of this knowledge. However, learning reusable skills from these logs requires addressing three challenges: (1) Goal Underspecification, since logs do not record the goal behind each action; (2) Non-Replayability, since past activity cannot be replayed to evaluate skill updates; and (3) Interleaved Trajectories, since logs may mix several tasks without marking their boundaries. To address these, we introduce TeleTune, a framework for learning a textual skill library from offline logs without recorded goals, cannot be replayed during optimization, and may interleave tasks. TeleTune uses action-prediction errors on logged trajectories to propose library edits and keep only those that improve held-out action-prediction accuracy, which we call skill-guided progress. The learned workflows also enable retrieval of demonstrations that cover the subgoals of a new task. At test time, the agent is provided with the learned library and the workflow-based retrieved demonstrations. Experiments on WorkArena and Online-Mind2Web show that TeleTune outperforms random retrieval, Agent Workflow Memory (AWM), and their combination. We find that the best baseline varies by setting, whereas TeleTune achieves average success rates of 77.1% and 80.6%, respectively, improving over the strongest baseline on each benchmark by 6.7% and 7.7%. Under the heaviest perturbation of the WorkArena training data,TeleTune keeps the highest average success rate at 68.5%, 6.3% above the strongest baseline. Our analyses show (1) skill optimization and workflow-based retrieval are complementary, (2) optimizing on fixed logs costs 5 to 75 times fewer tokens than validating the same edits with live episodes, (3) skill-guided progress tracks the live success rate.
Comments: Project Page: this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL)
Cite as: arXiv:2610.05437 [cs.AI]
  (or arXiv:2610.05437v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.05437

arXiv-issued DOI via DataCite

Submission history

From: Justin Chih-Yao Chen [view email]
[v1] Sun, 4 Oct 2026 18:19:33 UTC (693 KB)
[v2] Tue, 6 Oct 2026 03:17:00 UTC (693 KB)

来源:arXiv:cs.AI · arxiv.org