跳到正文

#视频

今日 17 条
今天10月1日周四
  1. HuggingFace Daily Papers(社区热门论文)39

    ThinkV2V:释放 MLLM 推理能力,实现指令引导的视频编辑

    ThinkV2V 是一个推理驱动的指令引导视频编辑框架,在视觉生成前显式激活 MLLM 思考,通过 MLLM-to-DiT 架构将思考转化为精炼的条件信号。它结合渐进式课程训练与推理时思考扩展,并构建了 ThinkV2V-150K 数据集和 ThinkV2V-Bench 基准。实验显示其 5B 规模 DiT 模型在复杂与标准编辑场景均达 SOTA,显著超越更大的 10B 规模基线。

  2. Arena.ai47

    Hidream-O1-Video-1.0 由 @HiDream_AI 刚刚登陆 Image-to-Video Arena,以 1456 分位列第 7! 该模型已在 @vivago_ai 上线,距 gemini-omni-flash 仅差 8 分,距 dreamina-seedance-2.0 和 2.5 均在 20 分以内。 恭喜 @HiDream_AI 发布!

    引用vivago.ai (HiDream)@vivago_ai

    New HiDream models just landed in vivago R1 Studio 🚀 Introducing: • HiDream-O1 Image 2.0 • HiDream-O1 Editing 1.5 • HiDream-O1 Video All three models are now available in vivago R1 Studio - bringing the latest HiDream image generation, editing, and video capabilities directly into your creative workflow. New models. New possibilities. Go make something the internet can’t ignore. 🔥

  3. MiniMax (official)48

    @HeyGen 发布 HeyGen Video,令人印象深刻!✨ 基于 MiniMax H3 构建,由 HeyGen 后训练,以更可及的成本为企业带来制作级视频。 很自豪能为这项工作提供基础,期待看到 HeyGen 团队将它带向多远。

    引用HeyGen@HeyGen

    We're releasing HeyGen Video, built for businesses that need production-quality video without production-level costs. Pricing starts at $0.01/s through October (50% off) Built on @Minimax_AI H3, post-trained by HeyGen. Learn more: https://developers.heygen.com/heygen-video-1.0-catalog

  4. MiniMax (official)43

    基于 MiniMax H3 构建,@Creatify_Labs 的 Boreal-H3 是一款专为广告优化的视频模型,在更精准遵循创意简报的同时,保持产品和角色的一致性。 很高兴看到 MiniMax H3 成为更多面向特定行业的前沿模型的基础!✨

    引用Creatify Labs@Creatify_Labs

    Introducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.

  5. 🚨 AI News | TestingCatalog46

    Creatify 推出广告视频模型 Boreal-H3 及由 Claude Opus 5.5 驱动的 Ad Agent,Boreal-H3 基于 MiniMax H3 用真实广告项目后训练。

    引用Creatify Labs@Creatify_Labs

    Introducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.

9月30日周三
  1. xAI:News(网页)49

    xAI 发布 Grok Imagine 1.5 Preview 图生视频模型,上线 xAI API

    xAI 最新图生视频模型 grok-imagine-video-1.5-preview 已通过 xAI API 开放预览。该模型可将单张静态图像转为流畅的电影感视频,支持最高 720p 分辨率,用户用自然语言提示词即可控制镜头运动、节奏与声音设计,并能保持原图的细节与光照。模型还支持逐帧生成并串联成更长场景,可用几行 Python 代码调用。

  2. Sarvam AI(网页)42

    Sarvam AI 推出 Sarvam Studio:面向多语言内容转换的智能体平台

    Sarvam AI 推出 Sarvam Studio,一个将语音、文本和文档工作流整合到同一工作空间的多语言内容转换平台,支持 AI 视频配音与智能体文档翻译。其配音在约 280 次与 ElevenLabs、YouTube Dub、Rask AI 的盲测对比中获得最高整体偏好,说话人相似度平均得分 0.88。该平台目前通过受控早期 beta 向政府、教育、媒体和出版领域的少量生产合作伙伴开放。

  3. Suno:Blog(网页)28

    Matt Steffanina 谈用 Suno 掌控舞蹈视频背后的音乐

    洛杉矶舞者、编舞师兼 DJ Matt Steffanina 在 Suno 博客访谈中表示,他用 Suno 为舞蹈内容创作原创音乐,把概念落地的时间从数天缩短到几分钟。他此前围绕他人音乐积累了数十亿播放量却不拥有底层资产,转向自制并拥有音乐后获得了更多机会与长期控制权。他近期在市中心地下通道拍摄的舞蹈视频即用 Suno 生成一首 house 曲目,成为其表现最好的 Hooks 之一。

  4. Suno:Blog(网页)33

    Dream Relic 如何用 Suno 为超现实视觉世界赋予声音

    AI 视觉艺术家 Dream Relic(Broc Vaughn)借助 Suno 的 Create 功能,把积压多年的歌词变成歌曲,用于 TikTok、Hooks、Spotify 及一张即将发行的全长专辑。一条深夜发布的视频引来数百条求歌名评论,让他重新重视音乐创作。他认为提示词与反复打磨等创作方向仍然关键,目标是让人感受作品而非关注工具。

  5. LMSYS:Blog(Chatbot Arena 团队)63

    SGLang-Diffusion 发布两个月:速度较初版提升至 2.5 倍,新增 ComfyUI 集成与 LoRA 支持

    LMSYS Chatbot Arena 团队复盘 SGLang-Diffusion 自 2025 年 11 月初发布两个月来的进展,当前版本(docker tag: lmsysorg/sglang:dev-pr-17247)比初版快最多 2.5 倍,在 NVIDIA GPU 上较其他方案最高快 5x。

    推荐理由:官方复盘发布两个月来的优化成果,速度较初版提升至 2.5 倍,并新增 ComfyUI 集成和 LoRA 支持,读者可评估是否迁移现有图像视频生成工作流。

  6. Baseten Base Labs:模型研究35

    Baseten 上线 Wan 2.2 T2V 文生视频 API

    Baseten 提供 Wan 2.2 T2V 文生视频 API,通过 POST 请求传入提示词即可生成视频,返回 base64 编码结果。支持 720x1280、1280x720、480x832、832x480、704x1280、1280x704、1024x704、704x1024 八种分辨率,可设置 sampling_steps、guide_scale 与 seed 参数。

  7. Every:最新文章(网页)47

    Vibe Check:GPT-6 Astra 能否成为独立电影人的突破?专业创意工作室实测

    创意工作室 Afterimage 用真实拍摄素材测试了 GPT-6 Astra 的视觉特效能力,包括替换移动镜头中的餐盘、在布鲁克林街景中添加建筑、让巨型游行花车出现在手持 iPhone 素材的窗外。部分镜头成功,部分失败,但团队最终开始构思此前不敢尝试的特效。此前 Opus 4.6 只能写确定性 Python 脚本做基础瑕疵与物体移除,Gemini Omni 虽能生成改动但缺乏精细控制。

  8. HuggingFace Daily Papers(社区热门论文)41

    WorldAttention:面向交互式视频世界模型的高效注意力架构

    研究者提出 WorldAttention,一种通过专用注意力内核与分层 KV cache 协同设计实现高效推理的注意力架构,用于文本条件交互式视频世界模型。其 Hybrid Sparse Attention 结合线性全局注意力与头自适应稀疏注意力,Hierarchical KV Cache 将历史 KV 对按语义索引分页存放于多级内存。

  9. HuggingFace Daily Papers(社区热门论文)37

    LongLive-Plug:面向视频生成的一次性蒸馏框架

    LongLive-Plug 是一个一次性蒸馏框架,将可复用能力以 LoRA 形式学习在基座模型上,实现免训练、即插即用地部署到兼容的下游模型。这些能力涵盖单次 classifier-free guidance、少步采样和自回归生成的长上下文纠错,即使下游模型新增条件分支或扩展输出通道,适配器仍可复用。

  10. HuggingFace Daily Papers(社区热门论文)37

    SoL-Refiner:一步精修实现 4K 高分辨率视频生成

    SoL-Refiner 是一种单步视频精修器,可将低分辨率模型输出经一次去噪直接转为 4K 视频,方法结合高分辨率持续训练、RL 后训练与最终一步蒸馏。在约 2K 分辨率下,其单步效果在 VBench 与 UniPercept 均值上超过所有外部精修器,在 3840×2176 下两项指标均优于三步的 LTX-2.3 Refiner。

  11. HuggingFace Daily Papers(社区热门论文)35

    APM-Bench:面向第一人称流式视频助手的跨会话持久记忆基准

    APM-Bench 将真实流式交互重构为多会话生活轨迹,包含 549 个会话、104 条轨迹和 2,719 个候选问题,覆盖客观题与开放题。评测显示现有方法难以同时兼顾可靠的长期召回、低开销与有效的主动协助,存在明显的效用—延迟—存储权衡。该基准还测试模型能否识别证据不足的情况。

  12. arXiv:cs.AI(全量分类)29

    AVIO:在音视频场景中学习添加与移除发声物体

    研究团队提出 AVIO,通过源条件特征调制改造预训练文本到音视频生成模型,让单一模型同时学会添加和移除发声物体,并支持仅凭指令编辑或可选视觉引导。配套数据集 AVIOBench 包含 37.9 小时配对音视频样本、覆盖 1,878 个目标物体名称,用共享身份与视觉掩码关联物体的视觉存在与声音贡献。参考帧课程训练逐步减少参考条件,使模型在添加物体时可通过可选参考控制外观与位置。

  13. arXiv:cs.AI(全量分类)35

    CoRe:协同演化奖励模型,缓解视频扩散模型中的潜在奖励攻击

    针对固定潜在奖励模型优化会导致"潜在奖励攻击"、生成视频感知与运动质量下降的问题,研究者提出协同演化奖励框架 CoRe,持续用生成器当前样本重训奖励模型并锚定真实视频偏好,防止生成器通过偏离数据分布刷高奖励。在 Wan2.1-T2V-1.3B 上,CoRe 的生成质量优于预训练模型和既有对齐方法,且避免了固定奖励优化的质量崩溃。