跳到正文
arXiv:cs.LG· Shunsuke Kamiya, Masanori Koyama, Seongcheol Jeong, Fumiya Uchiyama, Kenji Kubo, Kohei Hayashi, Masahiro Suzuki, Yutaka Matsuo·· 5 小时前AI 评分39

用读出反馈在推理时引导循环推理模型

Steering Recurrent Reasoners at Inference Time with Readout Feedback

AI 导读

研究者提出 Readout Feedback(RoFB),一种无需重训练的测试时干预方法,将循环模型的中间预测转化为 token 级成对耦合力注入潜在动态。在 AKOrN、ItrSA++、TRM 三个循环模型上测试 Sudoku 和 Maze,RoFB 在六组模型-任务组合中的四组取得明显提升、一组小幅提升,且四组正向结果达到仅靠增加步数或多轨迹采样无法实现的性能,计算成本相当或更低。

正文

View PDF HTML (experimental)

Abstract:Recurrent models, which repeatedly update latent states with shared computation blocks, have emerged as powerful architectures for solving complex reasoning tasks. Existing inference-time methods scale computation by running more steps or sampling more trajectories, but ignore information revealed within each trajectory. Here we show that recurrent models can be improved at inference time by using their own readout probabilities to steer latent dynamics without retraining. We introduce Readout Feedback (RoFB), a test-time intervention that converts intermediate predictions into token-wise pairwise coupling forces injected into the latent dynamics. Across three recurrent models (AKOrN, ItrSA++, TRM) on Sudoku and Maze, RoFB yields clear gains in four of six model-task pairs and a small gain in one. The four positive pairs achieve performance unattainable by merely running more steps or selecting from multiple trajectories, at comparable or lower computational cost. These results suggest that closed-loop steering of latent dynamics can serve as a complementary inference-time control mechanism for recurrent reasoning models.
Comments: 44 pages, 14 figures. Expanded analysis and experiments
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2608.24136 [cs.LG]
  (or arXiv:2608.24136v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2608.24136

arXiv-issued DOI via DataCite

Submission history

From: Shunsuke Kamiya [view email]
[v1] Tue, 25 Aug 2026 07:00:21 UTC (586 KB)
[v2] Fri, 2 Oct 2026 05:59:52 UTC (3,174 KB)

来源:arXiv:cs.LG · arxiv.org