跳到正文
The Decoder:AI News· Matthias Bastian·· 3 小时前AI 评分46

Reka AI 发布 19B 全能模型 Rho-1:单一模型统一处理文本、图像、视频与机器人控制

Reka AI's omni-model Rho-1 handles text, images, video, and robot control in a single model

AI 导读

Reka AI 发布 19B 参数全能模型 Rho-1 的研究预览版,在单一神经网络内处理并生成文本、图像、视频和机器人控制动作,所有模态以 token 形式在同一上下文窗口中运行,无需工具调用或外部模型。

正文

Reka AI has released a research preview of Rho-1. The 19-billion-parameter omni-model processes and generates text, images, video, and robot control actions in a single neural network. Unlike most AI systems that route tasks to specialized models, Rho-1 runs all modalities as tokens in one shared context window with no tool calls or external models. The model generates continuous video in real time and responds to new instructions on the fly without restarting.

The same weights that predict camera images also drive robot movements. To work around scarce robot training data, Reka AI built an inverse dynamics model that pulls control signals from ordinary internet videos. Rho-1 trained on 320 H100 GPUs over about three months.

Reka AI isn't new to multimodal AI. In April 2024, the company shipped Reka Core, a multimodal language model that competed with GPT-4, Claude 3, and Gemini Ultra on benchmarks. The release fits a broader push in AI research toward so-called world models.

AI News Without the Hype – Curated by Humans

Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.

来源:The Decoder:AI News · the-decoder.com