New HiDream models just landed in vivago R1 Studio 🚀 Introducing: • HiDream-O1 Image 2.0 • HiDream-O1 Editing 1.5 • HiDream-O1 Video All three models are now available in vivago R1 Studio - bringing the latest HiDream image generation, editing, and video capabilities directly into your creative workflow. New models. New possibilities. Go make something the internet can’t ignore. 🔥
#视频
#视频
今日 3 条
Arena.ai@arenaAI 评分4747引用vivago.ai (HiDream)@vivago_ai
MiniMax (official)@MiniMax_AIAI 评分4343引用Creatify Labs@Creatify_LabsIntroducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.
Artificial Analysis@ArtificialAnlysAI 评分5252
xAI:News(网页)AI 评分4949 xAI 发布 Grok Imagine 1.5 Preview 图生视频模型,上线 xAI API
xAI 最新图生视频模型 grok-imagine-video-1.5-preview 已通过 xAI API 开放预览。该模型可将单张静态图像转为流畅的电影感视频,支持最高 720p 分辨率,用户用自然语言提示词即可控制镜头运动、节奏与声音设计,并能保持原图的细节与光照。模型还支持逐帧生成并串联成更长场景,可用几行 Python 代码调用。
MiniMax:Blog(网页)精选AI 评分7171 MiniMax 开源 H3 通用视频生成模型,支持 2K 分辨率与原生立体声音频
MiniMax 正式开源下一代通用视频模型 MiniMax H3,支持文本、图像、视频、音频的多模态统一理解,可生成最长 15 秒、768p(经 H3-Regenerate-2K 可达 2K)、24 FPS 并带 32 kHz 立体声音频的视频。
推荐理由:官方开源公告给出了完整的三模块系统结构、模型规格和本地部署方式,读者可以据此评估在自己的视频生成流程中如何接入。
Meta AI:Blog(网页)AI 评分7474 Meta 发布 Muse Image 并预览 Muse Video
Meta Superintelligence Labs 发布 Muse Image 并预览 Muse Video,均为该实验室首批媒体生成模型。
Black Forest Labs:Blog(网页)精选AI 评分7171 Black Forest Labs 发布多模态基础模型 FLUX 3 并开放 Early Access
Black Forest Labs 发布多模态基础模型 FLUX 3,基于 Self-Flow 方法在统一架构内联合学习图像、视频和音频,现已开放 Early Access。
推荐理由:原文给出 FLUX 3 的多模态架构、早期对比评测数字和分阶段开放计划,读者可据此评估其对创作与具身场景的适用性。
Black Forest Labs:Blog(网页)精选AI 评分6666 Black Forest Labs 发布 FLUX 3 Video,文生视频与图生视频正式开放
Black Forest Labs 宣布 FLUX 3 Video 初版正式开放,可通过 BFL API 和部分合作伙伴使用。模型支持最长 20 秒、HD(720p)与 Full HD(1080p,经上采样)的视频生成,并伴随生成原生音频;能力包括文生视频、图生视频与关键帧、视频续接(最多 4 秒)、单次生成多镜头、多语言对白与唇形同步、Draft Mode 低成本预览等。
推荐理由:原文给出具体生成时长、分辨率、多语言与评测结果,读者可据此判断 FLUX 3 Video 在视频生成工作流中的可用性。
ModelScope@ModelScope2022AI 评分5858
Latent Space(RSS)AI 评分6262 Runway 推出 WorldPrompt:实时世界模型的工程实现
Runway 本月早些时候发布 GWM Worlds 2 研究预览,把高保真视频与音频生成变成实时交互式模拟,并提出新的输入格式 WorldPrompt,可固定环境部分要素(包括首帧)后生成带时间戳的事件序列,事件还能实时提示。
Google DeepMind:Blog(RSS)精选AI 评分6666 Google DeepMind 发布 Gemini 3.8 Live with Live Avatar
Google DeepMind 发布 Gemini 3.8 Live with Live Avatar,把近实时视频生成与语音对话模型结合,让对话 AI 具备动态视觉形象,支持精准唇形同步、自然表情和流畅轮次切换。
推荐理由:官方披露了实时视频与语音耦合的对话能力、异步工具调用和 97 种语言支持,可据此判断企业级数字人交互的落地边界。
Hao AI Lab@haoailabAI 评分4242
ViggleAI@ViggleAIAI 评分5353
MiniMax (official)@MiniMax_AIAI 评分3636引用SGLang@sgl_projectSGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀 On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup. No measured quality regression versus dense 50-step H3 across 103 test prompts. 🧵
MiniMax (official)@MiniMax_AIAI 评分5555
Hao AI Lab@haoailabAI 评分4646引用Physion Labs Official@Physion_Labs🐶🏁 Three models. One race. We independently tested H3 Max by @fal, minimax H3 by @MiniMax_AI, and FastH3 preview by @haoailab @haozhangml @wlsaidhi across robotics, animation, movies, and ads. H3 Max takes the overall lead 🏆. FastH3_preview, meanwhile, is a community OSS effort that’s already keeping up surprisingly well, even beating H3 on overall Prompt Adherence 🥳 The bigger gaps show up in Visual Integrity and Human Preference. Full evaluation: https://physionlabs.ai/blog/minimax-h3-evaluation
Fei-Fei Li@drfeifeiAI 评分5757引用World Labs@theworldlabsIntroducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
World Labs@theworldlabsAI 评分4242推出 Atlas: 全球首个多模态世界模型,可生成图像和视频帧,具备像素级精准的相机控制,并在 3D 中重建它们。 建模世界,移动相机,模拟空间与时间。

Hugging Face:Blog(RSS)AI 评分4848 NVIDIA Cosmos-H-Dreams:为手术机器人带来实时生成式仿真
NVIDIA 推出 Cosmos-H-Dreams,一个面向手术机器人的实时、动作条件生成式仿真器,通过 FlashDreams 推理库在单张 RTX PRO 6000 GPU 上运行,支持人与策略闭环交互控制。
美团 LongCat:HuggingFace 新模型AI 评分5353 美团 LongCat 发布开源模型 LongCat-Video-Avatar 1.5
美团 LongCat 团队发布开源音频驱动数字人视频生成框架 LongCat-Video-Avatar 1.5,基于 LongCat-Video 基础模型,支持 AT2V、ATI2V 和视频续写,兼容单路与多路音频输入。
美团 LongCat:HuggingFace 新模型AI 评分5555 美团 LongCat 发布开源音频驱动角色动画模型 LongCat-Video-Avatar
美团 LongCat 团队发布开源模型 LongCat-Video-Avatar,统一支持 Audio-Text-to-Video、Audio-Text-Image-to-Video 和 Video Continuation 三种生成模式,兼容单流与多流音频输入,提供单人与多人角色动画生成。
美团 LongCat:HuggingFace 新模型AI 评分6161 美团 LongCat 发布 13.6B 开源视频生成模型 LongCat-Video
美团 LongCat 团队发布 13.6B 参数的开源视频生成基础模型 LongCat-Video,统一支持文生视频、图生视频和视频续写三类任务。模型在 Video-Continuation 上预训练,可生成长达数分钟的视频而不出现色彩漂移或质量下降,并通过粗到细策略和 Block Sparse Attention 在数分钟内生成 720p、30fps 视频。
FireRedTeam:原创语音与多模态项目AI 评分3131 FireRedTeam 发布 DynamicPose:人体图像动画框架
FireRedTeam 推出 DynamicPose,一个用于人体图像动画的简单且鲁棒的框架。该框架以简洁稳健为设计目标,实现对人体图像的动画驱动。