跳到正文

#视频

今日 20 条
9月16日周三
  1. Hao AI Lab42

    在 ComfyUI 上体验 FastVideo 的 FastH3 V2 🚀🚀

    引用ComfyUI@ComfyUI

    FastH3 by FastVideo is now available in ComfyUI Video and native stereo audio, generated together, in seconds. Best for: → Previz and animatics that need lots of takes, quickly → Timing and dialogue tests before committing to a full-quality render → Short-form and social work on tight turnarounds Run it locally in ComfyUI and soon on Cloud ⬇️

9月15日周二
  1. MiniMax Design (H3)37

    一句简述 ➕ 一块画布 🟰 一整支制作团队,尽在 #MiniMaxDesign。

    引用Stefan 3D AI@Stefan_3D_AI

    I gave Astra one brief and it delivered the whole scene in Blender, camera direction included. All of it ran inside MiniMax Design. Astra sits there as the agent and talks to Blender over an official connector, so everything syncs straight onto the canvas. The final video is MiniMax H3, from the same canvas. #MiniMaxdesign - https://design.minimax.io/

  2. MiniMax (official)36

    H3 越来越快了。⚡️ @sgl_project + VDN-H3 现在让 MiniMax H3 在 8× B200 上突破 2 倍实时去噪——预热后端到端 9.0s 生成 14.4s 的 768p 视频,未测得质量下降。 开放模型通过开放生态持续进化。🚀

    引用SGLang@sgl_project

    SGLang-Diffusion with VDN-H3 now generates 14.4s of 768p video in just 9.0s 🚀 On 8× B200, 8 step denoising takes just 6.9s, reaching over 2× real time. The 9.0s figure covers the full generation request after warmup. No measured quality regression versus dense 50-step H3 across 103 test prompts. 🧵

9月14日周一
  1. MiniMax Design (H3)15

    最后召集 📢💥 释放你的怪兽,把大奖带回家👹💰

    引用Miora Design@Miora_Design

    Final call. Bestiary closes tonight, Sep 14 at 23:59 (PT). ⏳ $8,000 cash, 200,000 Credits, and ten Audience Choice awards are still on the table — and they go to the people who actually hit submit. One strange, beautiful short film, 30 seconds or longer, generated with MiniMax H3. Any myth, any era, any world you can dream up. The bestiary doesn't close itself. Finish your film before the clock runs out.

9月12日周六
  1. ViggleAI49

    你肯定没见过这个: GPT-6 Astra 用于场景与道具建模 + PINOC mcp 用于可动画的高斯泼溅角色

    引用PINOC@Viggle_PINOC

    GPT-6 Astra can now generate animatable Gaussian Splat characters. We connected it to the PINOC MCP and asked for a backrooms-style, Exit 8-ish game. We described the character we wanted and the motions. [MCP link in the comment 👇] Astra generated the character and every motion through PINOC through free preset animations and text to animation, and wrote the loop and the anomaly logic itself, and shipped the whole thing in a few sessions.

9月11日周五
9月10日周四
9月4日周五
9月3日周四
9月2日周三
  1. Hao AI Lab46

    很高兴看到 @Physion_Labs 对 FastH3 Preview、@MiniMax_AI 和 H3 Max 的独立评测 > FastH3 preview 作为社区开源成果,整体表现相当不错:它保持了很强的提示词遵循度,在这一维度上甚至超过了 H3…… 敬请期待我们下一版 FastH3,质量会更高!

    引用Physion Labs Official@Physion_Labs

    🐶🏁 Three models. One race. We independently tested H3 Max by @fal, minimax H3 by @MiniMax_AI, and FastH3 preview by @haoailab @haozhangml @wlsaidhi across robotics, animation, movies, and ads. H3 Max takes the overall lead 🏆. FastH3_preview, meanwhile, is a community OSS effort that’s already keeping up surprisingly well, even beating H3 on overall Prompt Adherence 🥳 The bigger gaps show up in Visual Integrity and Human Preference. Full evaluation: https://physionlabs.ai/blog/minimax-h3-evaluation

  2. Hao AI Lab49

    快来看 FastVideo 的 FastH3 Dense 在 @vllm_project vLLM-Omni 上部署运行!这就是开源的力量!我们必须赢! 感谢大家一直以来对 FastVideo 的支持与赞助!

    引用vLLM@vllm_project

    🎬Video generation faster than playback! 🚀MiniMax H3 on vLLM-Omni + FastVideo's FastH3: a complete 10.1s MP4 - video AND synchronized audio - rendered in 8.7s!⚡️ Thanks to @MiniMax_AI for the great Minimax H3 release, the FastVideo team @haoailab for open-sourcing FastH3 and helping on the serving integration, and @NVIDIAAI for the continued sponsorship and joint optimization efforts!

  3. Fei-Fei Li57

    Fei-Fei Li 宣布 World Labs 团队发布多模态世界模型 Atlas,称其从零训练、可生成像素级精准相机控制帧。Atlas 还能从单张输入图像重建大场景、通过重排视频帧模拟时空、从一张或多张图像原生输出 3D 空间,并将多张带位姿图像合成一致的 3D 世界,应用方向覆盖 VFX 到机器人。

    引用World Labs@theworldlabs

    Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.

  4. Google DeepMind:Blog(RSS)70

    Google DeepMind 为 Gemini 推出 agentic 视频理解功能

    Google DeepMind 推出 agentic video understanding,覆盖 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite,通过智能体循环动态调用原生视频工具按需检索画面、音频和字幕,而非固定帧率静态处理。

    推荐理由:原文给出了具体降本增效数字、适用模型和接入方式,开发者可据此评估是否切换视频分析流程。

9月1日周二
8月31日周一
  1. Hao AI Lab52

    Reactor 宣布与 Hao AI Lab 合作,基于 FastVideo 的 FastH3 上线无限直播流,入口为 https://twitch.tv/dereactor,并宣布将全部开源。作者转发称赞社区进展迅速,强调项目完全开源。

    引用reactor@reactorworld

    At Reactor we ❤️ open-source. We teamed up with @haoailab to ship an infinite live stream powered by FastVideo’s FastH3: https://twitch.tv/dereactor And we’re open-sourcing everything!

8月28日周五
8月27日周四
  1. Saining Xie43

    很高兴看到 RAE 扩展到视频!

    引用Minghui Guo@MinghuiGuo77

    🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE

8月21日周五
8月17日周一
7月27日周一
7月20日周一
7月3日周五
7月1日周三
6月3日周三
  1. Saining Xie42

    大脑如何从(可能不完整且有噪声的)视觉观察中构建并追踪世界的内部状态? 我相信视觉状态追踪将成为未来几年视觉领域的重大挑战,我希望这个基准能成为一个有用的起点。enjoy!

    引用Sihyun Yu@sihyun_yu

    Can MLLMs actually track what's happening in a video? Introducing VSTAT 🎯, our new benchmark for visual state tracking. The tasks are simple: count cups, read typed words, count page flips. Humans solve them easily. MLLMs don't. https://vision-x-nyu.github.io/vstat-site/ 🧵 [1/11]

5月22日周五
  1. 美团 LongCat:HuggingFace 新模型30

    美团 LongCat 发布 WBench-weights:交互式视频世界模型评测权重

    美团 LongCat 在 HuggingFace 发布 WBench-weights,整合了 WBench 交互式视频世界模型评测框架所需的模型权重,方便社区直接部署。WBench 从视频质量、设定遵循、交互遵循、一致性和物理合规五个维度评测世界模型,包含 289 个测试用例和 1,058 轮交互。权重仅限学术研究与评测用途,可通过 huggingface-cli 下载。

5月21日周四
5月18日周一
  1. Google DeepMind:Blog(RSS)80

    Google DeepMind 发布 Gemini Omni Flash,支持多模态输入生成与对话式编辑视频

    Google DeepMind 发布 Gemini 家族首个模型 Gemini Omni Flash,可将图像、音频、视频和文本组合作为输入,生成高质量视频,并支持通过自然语言多轮编辑视频、保持角色和场景一致。

    推荐理由:原文介绍了 Gemini Omni Flash 的多模态输入、对话式视频编辑能力和开放范围,读者可以判断它对视频创作流程的影响。

4月12日周日
3月12日周四
  1. OpenAI Developers(RSS)67

    OpenAI 发布 Sora 2 提示词指南

    OpenAI 发布 Sora 2 提示词指南,更新至最新 API 能力,包括角色引用(可上传动物或对象并复用)、1920×1080 或 1080×1920 高分辨率导出、时长上限从 12 秒提高到 20 秒、基于完整原始片段的视频续写,以及支持异步批量生成的 Batch API。

    推荐理由:OpenAI 官方系统讲解 Sora 2 提示词写法,并覆盖角色引用、更长时长和视频续写等新 API 能力。

2月7日周六
  1. FireRedTeam:原创语音与多模态项目45

    FireRedTeam 发布 FireRed-OpenStoryline AI 视频剪辑智能体

    FireRedTeam 推出 AI 视频剪辑智能体 FireRed-OpenStoryline,通过自然语言交互、LLM 驱动的规划与精准工具编排,将手动剪辑转变为意图驱动的导演式创作。该智能体支持人在回路中的透明创作流程,并提供可复用的 Style Skills,以保持叙事风格的一致与专业。

12月13日周六
10月25日周六
  1. 美团 LongCat:HuggingFace 新模型61

    美团 LongCat 发布 13.6B 开源视频生成模型 LongCat-Video

    美团 LongCat 团队发布 13.6B 参数的开源视频生成基础模型 LongCat-Video,统一支持文生视频、图生视频和视频续写三类任务。模型在 Video-Continuation 上预训练,可生成长达数分钟的视频而不出现色彩漂移或质量下降,并通过粗到细策略和 Block Sparse Attention 在数分钟内生成 720p、30fps 视频。