arXiv:cs.LG· Antoine Edy, Max Conti, Victor Xing, Marc-Antoine Allard, Nawfal Benhamdane, Gautier Viaud·· 6 小时前AI 评分52
DAEDALUS:从自生成任务自举智能体记忆
DAEDALUS: Bootstrapping Agent Memory from Self-Generated Tasks
AI 导读
论文提出 DAEDALUS,通过 explorer 与 solver 双智能体从自生成练习中自举可复用的智能体记忆,无需已有任务或 oracle 验证器。
正文
Abstract:LLM agents often lack the operational knowledge to act reliably in new environments, as they must discover specific tool behaviors or environment conventions on their own. Without memory of past attempts, they repeat the same mistakes across tasks, leading to more task failures and longer trajectories. To address this, agentic systems typically rely on human-written guidelines or on procedural memory built from training tasks and an oracle verifier, both of which require prior knowledge of the environment. We present DAEDALUS, a method for bootstrapping reusable agent memory from self-generated practice without existing tasks or oracle verifiers. DAEDALUS pairs two agents: an explorer that interacts with the environment to generate challenging yet solvable tasks, and a solver that attempts them. A heuristic is derived from each solver failure and accepted only after the solver repeatedly succeeds with that heuristic in context. These outcomes also provide feedback for the explorer to refine the difficulty of future tasks. Accepted heuristics are then consolidated into a memory bank for test-time use. Across AppWorld, $\tau^2$-bench, and AutomationBench, DAEDALUS improves mean success rates by up to 15.9 points and pass^5 by up to 2.2x over a no-memory baseline, and is competitive with methods using training tasks, at a lower inference cost than most. We show that performance gains already emerge with a small exploration budget, and that its heuristics also benefit agents from other model families. Our ablations further reveal that solver traces provide the key information needed to derive effective heuristics, while factorizing early discoveries makes exploration more cost-efficient. Beyond memory construction, we find that the tasks generated by DAEDALUS can serve as a proxy for benchmark tasks when ranking models by performance. Code and artifacts: this http URL.
| Comments: | 9 pages (31 including Appendix), 8 figures (11 including Appendix). We release the code and artifacts, including generation and inference traces, at this https URL |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.08048 [cs.AI] |
| (or arXiv:2610.08048v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08048 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Antoine Edy [view email]
[v1]
Tue, 6 Oct 2026 09:46:32 UTC (751 KB)
来源:arXiv:cs.LG · arxiv.org