跳到正文
原文
arXiv:cs.AI(全量分类)· Xisen Jin, Jingheng Li, Zhenglun Chen, Junyi Du, Xiang Ren·· 5 小时前AI 评分44

ReLiveGym:评估在数周重放现实中长期运行的 LLM 智能体

ReLiveGym: Evaluating Long-Lived Agents over Weeks of Replayed Reality

AI 导读

研究者推出 ReLiveGym,一个针对长期运行 LLM 智能体的诊断评估环境,让智能体在按时间顺序重放的数周真实新闻、市场和社交媒体流中稀疏行动。团队在八种基础语言模型上测试了模型选择与 harness 设计对长期任务表现的影响,发现"何时行动"是关键的 harness 设计维度,且最优设计因任务和模型而异。研究还评估了基于事后反馈的持续学习对性能与失败模式的作用。

正文

View PDF

Abstract:As large language model (LLM) agents become widely adopted, they are increasingly deployed for tasks that require persistent monitoring or recurring actions (e.g., market analysis). These agents are expected to operate unattended for days or weeks, act at the right timing, and adapt to the dynamic environment over time. These challenges are not fully captured in the existing long-horizon agent work, as they often consider a static environment that is not temporally changing. We introduce ReLiveGym, a diagnostic evaluation environment of long-lived tasks in which agents act sparsely over simulated weeks of chronologically replayed real-world news, market, and social-media streams. The tasks span diverse levels of time sensitivity, reasoning intensity, and recurrence. Across eight base language models, we investigate how model choice and harness design affect agent performance on such long-lived tasks. Our results show that how agents determine when to act arises as an important harness-design axis for long-lived tasks; and that the optimal design varies across tasks and sometimes model choices as well. We also evaluate how continuous learning from hindsight feedback affects performance and addresses failure modes observed in these long-lived tasks. These findings indicate model choice, action timing mechanism, and use of feedback as important considerations in the design of long-lived agents. Code: this https URL
Comments: 9 pages. Preprint
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.00710 [cs.AI]
  (or arXiv:2610.00710v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.00710

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xisen Jin [view email]
[v1] Wed, 30 Sep 2026 20:59:33 UTC (594 KB)

来源:arXiv:cs.AI(全量分类) · arxiv.org