arXiv:cs.AI· Renxi Wang, Rifo Ahmad Genadi, Bilal El Bouardi, Yongxin Wang, Fajri Koto, Zhengzhong Liu, Timothy Baldwin, Haonan Li·· 3 小时前
AgentFly:用统一资源系统扩展智能体强化学习
AgentFly: Scaling Agentic Reinforcement Learning with Unified Resource System
AI 导读
AgentFly 是一个智能体强化学习框架,通过统一资源层把沙盒、模型服务和外部 API 视为统一调度的类型化资源,支持按工具获取、多轮复用和异步背压。框架采用智能体、rollout、上下文和底层资源四层设计,并提供预置工具与环境,首次给出跨框架吞吐量对比。
正文
Abstract:Methods to build LLM agents have evolved from prompt engineering and supervised finetuning to agentic reinforcement learning (agentic RL). However, agentic RL remains bottlenecked by its surrounding systems: agents must interact with heterogeneous environments, such as sandboxes, model services, and external APIs. Their allocation, reuse, and lifecycle dominate rollout cost and cap the scale at which training becomes practical. In this work, we present AgentFly, an agentic RL framework built with a unified resource layer that treats each of these environments as a distinct, typed resource scheduled through one engine, with per-tool acquisition for multi-turn reuse, asynchronous backpressure, and rollout versus global-scoped lifecycles. AgentFly adopts a four-layer design: (I) agent layer that abstracts the agent, tool, and reward concepts, decomposing agentic RL into defining agents, tools, and reward functions; (II) rollout layer that composes these into agent loops and computes rewards; (III) context layer that organizes rollouts, injects contextual information, and arranges resources; and (IV) a low-level resource layer that performs resource management. We provide a suite of prebuilt tools and environments, demonstrate successful agent training across multiple tasks and models, and report the first controlled cross-framework throughput comparison against agentic RL frameworks.
| Subjects: | Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.5 |
| Cite as: | arXiv:2507.14897 [cs.AI] |
| (or arXiv:2507.14897v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2507.14897 arXiv-issued DOI via DataCite |
Submission history
From: Renxi Wang [view email]
[v1]
Sun, 20 Jul 2025 10:22:36 UTC (947 KB)
[v2]
Thu, 8 Oct 2026 17:02:18 UTC (2,847 KB)
来源:arXiv:cs.AI · arxiv.org