Fuli Luo· @_LuoFuli · X·· 15 天前AI 评分38
AI 导读
小米 MiMo-V2.6 正处于 RL 训练中,团队从三方面扩展规模:计算(每步约 2B tokens、1568 prompts × 16 rollouts、全异步)、环境与 harness(单次运行混合多 harness 的多任务 agentic RL)、评分算力(组内 agentic 信用分配,结合测试用例与 rubric 奖励)。
正文
Nearly half a year of silence. We spent it studying one problem: how far RL can scale.
MiMo-V2.6 is in the middle of its RL run right now. Three things we scaled: compute (~2B tokens per step, 1568 prompts × 16 rollouts, fully async), environments and harnesses (multi-task agentic RL, mixed across multiple harnesses in one run), and grader compute (agentic in-group credit assignment, with test-case and rubric-based rewards). We'll open-source the details piece by piece over the coming weeks.
Streaming the run: https://mimo.xiaomi.com/rl/
来源:Fuli Luo · x.com