More than 25 million people are using Google Flow every month to dream up new ideas, create stories, and build cool things. Thank you. We’re continuing the 50 additional daily credits for all users. Dive in and keep creating.
#视频
#视频
今日 4 条
Josh Woodward@joshwoodwardAI 评分3939引用Google Flow@FlowbyGoogle
OpenAI:官网动态(RSS · 排除企业/客户案例)AI 评分2323 Higgsfield AI 借助 GPT-6 Astra 一天内上线新视频功能
Higgsfield AI 借助 GPT-6 Astra 在一天内上线新视频功能,让小型企业更轻松地制作视频广告,并更快将新创意工具推向市场。
MiniMax Design (H3)@Hailuo_AIAI 评分3636🔥社区从不停下折腾的脚步。 不只是基于 H3 做开发,还在不断深入内部,寻找让它更聪明的新方法。
引用Kamimoto(かみもと)@sep_is_heim流行のJevをMiniMax H3に組み込んで、動画生成を高速化してみた!Attention処理のスパース化にJevを使用。 ・層ごとにJevが重要度を判定(4step 49層が対象) ・Jevがスパース率1%, 3%, 5%, 10%を選択 RTX4070で6分7秒→3分34秒で41.7%短縮!動画生成中にJevクラウドに問合せしているのに速い!
MiniMax Design (H3)@Hailuo_AIAI 评分5454引用852話(hakoniwa)@8co28この3x3画像を1枚 r2v でAI動画化すると以下の設定でおおよそそのまま映像になる Minimax H3 Max r2v 480p 15秒 Prompt調整:Quality 参照強度:標準 「左上から右下のパネルにカットが切り替わる2Dアニメーション、 複数パネル禁止、BGM禁止」
MiniMax (official)@MiniMax_AIAI 评分4040引用Nunchux AI@NunchuxAIIntroducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.
vLLM 官方博客(RSS)AI 评分4444 vLLM 集成 PyNvVideoCodec:用 GPU 硬件解码扩展多 GPU 视频描述吞吐
vLLM 集成 PyNvVideoCodec(NVDEC 的 Python 接口),把视频解码从 CPU 转移到 NVIDIA GPU,消除多 GPU 视频描述任务的 CPU 瓶颈。
MiniMax Design (H3)@Hailuo_AIAI 评分3737一句简述 ➕ 一块画布 🟰 一整支制作团队,尽在 #MiniMaxDesign。
引用Stefan 3D AI@Stefan_3D_AII gave Astra one brief and it delivered the whole scene in Blender, camera direction included. All of it ran inside MiniMax Design. Astra sits there as the agent and talks to Blender over an official connector, so everything syncs straight onto the canvas. The final video is MiniMax H3, from the same canvas. #MiniMaxdesign - https://design.minimax.io/
MiniMax Design (H3)@Hailuo_AIAI 评分2020引用Koldo Huici@koldo2kI used the work generated with GPT-6 Astra to build out a set, a story. Now I'm bringing it to life with @Hailuo_AI Minimax H3. Here's a tip that's worked well for me 👇
MiniMax (official)@MiniMax_AIAI 评分3636引用SGLang@sgl_projectSGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀 On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup. No measured quality regression versus dense 50-step H3 across 103 test prompts. 🧵
MiniMax Design (H3)@Hailuo_AIAI 评分1515引用Miora Design@Miora_DesignFinal call. Bestiary closes tonight, Sep 14 at 23:59 (PT). ⏳ $8,000 cash, 200,000 Credits, and ten Audience Choice awards are still on the table — and they go to the people who actually hit submit. One strange, beautiful short film, 30 seconds or longer, generated with MiniMax H3. Any myth, any era, any world you can dream up. The bestiary doesn't close itself. Finish your film before the clock runs out.
MiniMax (official)@MiniMax_AIAI 评分5555
Google DeepMind:Blog(RSS)精选AI 评分7070 Google DeepMind 为 Gemini 推出 agentic 视频理解功能
Google DeepMind 推出 agentic video understanding,覆盖 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite,通过智能体循环动态调用原生视频工具按需检索画面、音频和字幕,而非固定帧率静态处理。
推荐理由:原文给出了具体降本增效数字、适用模型和接入方式,开发者可据此评估是否切换视频分析流程。
vLLM 官方博客(RSS)AI 评分5252 vLLM-Omni 详解 MiniMax H3 生产级服务:系统优化与 FastH3 实时生成
vLLM-Omni 团队发布 MiniMax H3 服务优化详解,先对完整管线做无损系统级优化,相比 Diffusers 将完整响应延迟降低 30.8%(1.445x)。
Google DeepMind:Blog(RSS)精选AI 评分6262 Google DeepMind 发布 Gemini Omni 1.1 Flash,强化生成式视频控制能力
Google DeepMind 发布 Gemini Omni 1.1 Flash,通过 Gemini API 和 Google AI Studio 面向开发者提供新的创意控制与生成式视频能力。
推荐理由:原文来自官方,列出了各能力具体参数、速度和价格差异,开发者可据此评估是否接入自己的视频工作流。
Saining Xie@sainingxieAI 评分4343引用Minghui Guo@MinghuiGuo77🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE
vLLM 官方博客(RSS)AI 评分5050 vLLM-Omni 推出分布式逐层卸载,可在多卡上高效运行 200B+ 级 DiT 模型
vLLM-Omni 团队发布 Distributed Layerwise Offload,让超过单卡 HBM 的视频生成模型(如 Cosmos3-Super 64B / 124 GB)可在多块 NPU 或 GPU 上运行。
Hugging Face:Blog(RSS)AI 评分4848 NVIDIA Cosmos-H-Dreams:为手术机器人带来实时生成式仿真
NVIDIA 推出 Cosmos-H-Dreams,一个面向手术机器人的实时、动作条件生成式仿真器,通过 FlashDreams 推理库在单张 RTX PRO 6000 GPU 上运行,支持人与策略闭环交互控制。
Google DeepMind:Blog(RSS)AI 评分5252 Google DeepMind 与 A24 宣布研究合作,Google 同时投资 A24
Google DeepMind 与 A24 宣布一项研究导向的合作,让 A24 及其电影创作者参与塑造新工具和工作流,Google DeepMind 则获得艺术家的反馈与指导,合作将跨多个项目长期进行。此外,Google 已对 A24 进行投资。
Google DeepMind:Blog(RSS)精选AI 评分6868 Google DeepMind 发布 Nano Banana 2 Lite 并向开发者开放 Gemini Omni Flash
Google DeepMind 发布 Nano Banana 2 Lite(gemini-3.1-flash-lite-image),称其为 Nano Banana 家族中最快。
推荐理由:官方同时推出 Nano Banana 2 Lite 与 Gemini Omni Flash,给出了价格、延迟和限制等可核对的发布细节。
Saining Xie@sainingxieAI 评分4242大脑如何从(可能不完整且有噪声的)视觉观察中构建并追踪世界的内部状态? 我相信视觉状态追踪将成为未来几年视觉领域的重大挑战,我希望这个基准能成为一个有用的起点。enjoy!
引用Sihyun Yu@sihyun_yuCan MLLMs actually track what's happening in a video? Introducing VSTAT 🎯, our new benchmark for visual state tracking. The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't. https://vision-x-nyu.github.io/vstat-site/ 🧵 [1/11]
Google DeepMind:Blog(RSS)精选AI 评分8080 Google DeepMind 发布 Gemini Omni Flash,支持多模态输入生成与对话式编辑视频
Google DeepMind 发布 Gemini 家族首个模型 Gemini Omni Flash,可将图像、音频、视频和文本组合作为输入,生成高质量视频,并支持通过自然语言多轮编辑视频、保持角色和场景一致。
推荐理由:原文介绍了 Gemini Omni Flash 的多模态输入、对话式视频编辑能力和开放范围,读者可以判断它对视频创作流程的影响。
OpenAI Developers(RSS)精选AI 评分6767 OpenAI 发布 Sora 2 提示词指南
OpenAI 发布 Sora 2 提示词指南,更新至最新 API 能力,包括角色引用(可上传动物或对象并复用)、1920×1080 或 1080×1920 高分辨率导出、时长上限从 12 秒提高到 20 秒、基于完整原始片段的视频续写,以及支持异步批量生成的 Batch API。
推荐理由:OpenAI 官方系统讲解 Sora 2 提示词写法,并覆盖角色引用、更长时长和视频续写等新 API 能力。
OpenAI Developers(RSS)AI 评分2323 Sora、ImageGen 与 Codex:创意生产的下一波浪潮
Sora、ImageGen 与 Codex 被放在同一条多模态流水线中,串联视频、图像与代码生成工具,指向创意生产工作流。内容聚焦这些工具之间的衔接方式,而非单一模型的独立能力。
OpenAI Developers(RSS)AI 评分4444 OpenAI 发布 Sora starter app:基于 Sora Video API 的 NextJS 示例应用
OpenAI 推出基于 Sora Video API 和 OpenAI SDK 的 NextJS 示例应用 Sora starter app,提供文本提示词加可选图像输入来生成和 remix 视频的简易 UI。