跳到正文
原文
Arena.ai· @arena · X·· 2 小时前精选AI 评分80
AI 导读

Arena 宣布小米 MiMo-V2.6-Pro 和 MiMo-V2.6-Flash 上线 Agent Arena。Pro 在 8.1K+ 真实智能体会话中净提升 +3.17%,列开源模型第 5,较 MiMo-V2.5-Pro(第 13,-7.23%)提升 9 位和 10.4 个百分点;其 Confirmed Success 得分 +7.35%,列开源模型第 2。

推荐理由

Arena 官方公布了 MiMo-V2.6 两款模型在真实智能体会话中的排名、净提升幅度和成本数据,可与引用的官方发布对照阅读。

正文 · AI 翻译

MiMo-V2.6-Pro 和 MiMo-V2.6-Flash 由 @XiaomiMiMo 打造,现已登陆 Agent Arena!

MiMo-V2.6-Pro 在开源模型中排名第 5,在 8.1K+ 真实世界智能体会话中净提升 +3.17%。此次发布相比 MiMo-V2.5-Pro 提升了 +10.4 个百分点,排名上升 9 位——后者排名第 13,净提升为 -7.23%!

其最强信号是 Confirmed Success(用户明确反馈任务成功完成),在此项上得分 +7.35%,在开源模型中排名第 2!

MiMo-V2.6-Flash 在开源模型中排名第 9,在 13K+ 真实世界智能体会话中净提升 -0.57%。

其每任务中位成本为 $0.04(比 MiMo-V2.6-Pro 低 56%),同时也登上了 Agent Arena 帕累托前沿!见下方其位置。

恭喜 @XiaomiMiMo 团队发布!

引用Xiaomi MiMo@XiaomiMiMo
隆重推出 Xiaomi MiMo-V2.6 —— Pro 与 Flash。 前沿智能,全模态,公开构建。 🔹 两款全模态模型,通过规模化强化学习不断进阶 🔹 Pro 在大多数 agent 基准测试中表现与 Claude Opus 5 和 GPT-5.6 Sol 相当 🔹 Pro 在 Artificial Analysis Intelligence Index 上得分 46 —— 开源模型中最高的 🔹 更强的编程、计算机使用、3D 推理和创作能力 🔹 开放模型权重、技术报告、RL 环境和训练代码 Blog:https://mimo.xiaomi.com/mimo-v2-6
原文

Introducing Xiaomi MiMo-V2.6 — Pro & Flash. Frontier intelligence, all the modalities, built in public. 🔹 Two omnimodal models, advancing through scaled reinforcement learning 🔹 Pro performs on par with Claude Opus 5 and GPT-5.6 Sol across most agent benchmarks 🔹 Pro scores 46 on the Artificial Analysis Intelligence Index — the highest among open-source models 🔹 Stronger coding, computer use, 3D reasoning and creative capabilities 🔹 Open model weights, technical report, RL environments and training code Blog:https://mimo.xiaomi.com/mimo-v2-6

在 X 查看被引用的帖子

来源:Arena.ai · x.com