NVIDIA AI· @NVIDIAAI · X·· 9 小时前AI 评分48
AI 导读
CoreWeave 新服务借助 NVIDIA Dynamo 的 ModelExpress 和 Router,在 RL 后训练中实现比基线快 15 倍的模型重载,同时将停机时间降至最低。
正文
Congrats @CoreWeave on RL Rollouts!
RL post-training involves a lot of back and forth: train the model, generate responses, then train again. Inference workers need to load the updated model weights each time. As models get bigger, that can leave GPUs waiting.
CoreWeave’s new service uses ModelExpress and Router in NVIDIA Dynamo to speed up those reloads with minimal downtime.
Working with us and @youdotcom, CoreWeave achieved 15× faster model reloads compared with its baseline while post-training Nemotron 3.5 Lightning.
Check out their blog below for details
ICYMI: CoreWeave Forge is here 🎉 A production trace that never reaches the next training run is a signal you paid for and threw away. Most teams do it every day, because the tool that catches the trace and the tool that runs the training came from different vendors and were never built to talk. Forge closes that gap by unifying @wandb, post-training from @OpenPipeAI, and @marimo_io notebooks in one connected environment with CoreWeave Training, Inference, Sandboxes, and Registry. Run, observe, curate, improve, evaluate. The traces you flag in production become the datasets you train on and the evaluations you gate with. Open across any model, framework, or cloud. @MasterClass and @canva are already building on it. Free, Pro, and Enterprise available today: https://crwv.co/utcAw在 X 查看被引用的帖子
来源:NVIDIA AI · x.com