跳到正文
arXiv:cs.AI· Ankur Samanta, Yonathan Efroni, Paul Sajda, Kaveh Hassani, Anirudh Goyal·· 5 小时前AI 评分41

MIRA:面向长时程研究智能体的元推理架构

Learning What to Investigate Next: Meta-Reasoning for Long-Horizon Research Agents

AI 导读

研究者提出 MIRA,一种将研究资源分配与执行分离的分层架构:外层元推理器从持久研究记录中整理上下文并生成下一步调查的工作指令,内层执行器负责执行,使执行成为元推理动作之间的状态转移。

正文

View PDF HTML (experimental)

Abstract:Long-horizon research agents must decide both how to investigate and what to investigate next as evidence accumulates. This is hard to learn because such decisions are sparse in long execution traces, and their consequences may emerge several investigations later. We introduce Meta-reasoning for Iterative Research Agents (MIRA), a hierarchical architecture separating research allocation from execution. An outer-loop meta-reasoner curates context from a persistent research record, then writes a work order for the next investigation or ends the episode. A fresh inner-loop executor carries out each work order, making execution part of the transition between meta-reasoning actions. Without policy training, MIRA improves long-horizon inference and allocates additional compute more effectively in theorem proving and open-ended neural-architecture research. Its decision boundaries also provide natural units for credit assignment. At each boundary, we train a generative critic to forecast expected remaining return from partial states, outperforming token-level alternatives. Cross-environment pretraining improves forecasting and adaptation, yielding a transferable prior for valuing partial progress. We use this prior to initialize MIRA-AC, a generative actor-critic jointly trained to forecast remaining return and choose the next investigation, without a separate critic model. MIRA-AC concentrates policy optimization on meta-reasoning decisions, enabling efficient long-horizon reinforcement learning without directly optimizing the longer execution traces they initiate. Training MIRA-AC on the model's own proxy hill-climbing signals improves gold performance across four autoresearch environments; the actor transfers with cross-environment value initialization. Together, these results show that meta-reasoning can be learned as an explicit policy for directing long-horizon autonomous research.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.02525 [cs.AI]
  (or arXiv:2610.02525v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.02525

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Ankur Samanta [view email]
[v1] Thu, 1 Oct 2026 21:54:47 UTC (1,773 KB)

来源:arXiv:cs.AI · arxiv.org