跳到正文
原文
Prime Intellect(网页)·· 13 小时前AI 评分53

Prime Intellect 提出 ECHO:在 RL 中加入世界模型训练提升智能体泛化

ResearchJUN 05TH, 2026True Agents Model the World

AI 导读

Prime Intellect 发布研究文章,提出在 RL 训练中同时用 SFT(正恒定优势、无额外前向开销)让模型预测工具输出,即 ECHO 方法。

来源:Prime Intellect(网页) · primeintellect.ai