跳到正文

#视频

今日 19 条
9月28日周一
9月27日周日
9月26日周六
9月25日周五
  1. Google Research67

    Google Research 发布 AI 视频联合导演框架,实现连贯长视频生成

    Google Research 发布 AI video co-director,一个构建在 Gemini 和 Veo 之上的多智能体编排框架,用于生成连贯的多镜头长视频,并原生继承 SynthID 水印等安全机制。

    推荐理由:Google 把长视频生成拆成四个可组合的智能体框架,并给出三套自建基准与量化结果,便于对照现有方案。

  2. Google DeepMind:Blog(RSS)66

    Google DeepMind 发布 Gemini 3.8 Live with Live Avatar

    Google DeepMind 发布 Gemini 3.8 Live with Live Avatar,把近实时视频生成与语音对话模型结合,让对话 AI 具备动态视觉形象,支持精准唇形同步、自然表情和流畅轮次切换。

    推荐理由:官方披露了实时视频与语音耦合的对话能力、异步工具调用和 97 种语言支持,可据此判断企业级数字人交互的落地边界。

9月24日周四
  1. MiniMax (official)29

    更多用 MiniMax-H3 构建的方式。🐮 很高兴看到模型优化与 AMD 推理工程结合,为创作者带来更快的反馈循环。⚡️

    引用Nunchux AI@NunchuxAI

    5s of video, generated in 1.3s! Nunchux brings @MiniMax’s MiniMax-H3 to @AMD MI355X with up to 26.7× faster inference than SGLang on 8 GPUs @AMDServer And with streaming generation, you can change the prompt as the video plays, steering what happens next. One stack, optimized across hardware. Free access to MiniMax-H3 is coming soon! Join the waitlist now at http://nunchux.ai. Our blog: https://www.nunchux.ai/blog/video-generation-on-amd-mi355x #AI #VideoGeneration #MiniMaxH3

  2. Sherwin Wu48

    听完 @AcquiredFM 的迪士尼那期节目后,再读 Jeffrey Katzenberg 这篇帖子,感受完全不同。 想不到还有谁比他更适合在 AI 时代,就创造力的未来分享一份乐观的看法!

    引用Jeffrey Katzenberg@jeffreykWNDR

    The World is Changing: AI For Creativity By Jeffrey Katzenberg A few months ago, I sat in my office in Silicon Valley and watched as a tech founder showed me something extraordinary. On the screen was a fully realized, beautifully lit, well-composed animated scene. It was stunning and it made me feel exactly what I felt in 1986 watching Luxo Jr. That was the first time I watched a computer-animated 3D character take a breath and seem, against all reason, to have life. It left me in awe. Later that day, I received a text from an artist I've known for thirty years, 350 miles to the south, in the city where I spent most of my career. After seeing a similar video, she texted: "Is this the end of us?" My answer was, "Certainly not.” I have spent the better part of the last decade in Silicon Valley, but the heart of my career has been in Hollywood. Being deeply connected to both worlds means I have deep loyalties to each and a responsibility to speak honestly to both. In 2023, I said that these new AI tools would cut the time and cost of producing world-class animation by as much as ninety percent within three years. Some colleagues were alarmed, many were furious. There is growing fear and resistance surrounding AI within the creative community. I deeply understand it, because I've spent countless hours walking through animation studios watching gifted artists bent over their desks, rebuilding a single second of film for the tenth time because the ninth version wasn't quite right. I've sat in screening rooms where four years of people's labor played out in minutes, and I knew the name of every person that had spent countless hours bringing those images to life. The creative process is a calling, there's really no other way to describe it. From the outside some see resistance. From the inside, it is love. People do not fight this hard for things they don't care about. The pushback coming out of Hollywood represents the collective effort of people who are deeply passionate about their craft. Is History Repeating Itself? The history here is more complicated than either side may realize. In 1906, the most famous composer in America, John Philip Sousa, published an essay titled “The Menace of Mechanical Music." He warned that the phonograph would become "a substitute for human skill, intelligence and soul." Sousa's fight was not really about the machine, it was about money. The machines were playing his compositions, and the men who built them weren't paying him a cent. His campaign helped create the Copyright Act of 1909. He did not stop the technology. He changed the terms under which it could use his work. A hundred years ago, sound came to the movies. We remember it now as a miracle, and it was. What we forget is who paid for it. Before sound, tens of thousands of musicians made their living in the orchestra pits of movie houses, scoring every film live, every night, in towns all over the world. When the soundtrack arrived, the work of one composer and one orchestra was recorded for a film that went into thousands of theaters. The union fought back with everything it had, taking out newspaper ads across the country warning against the menace of "canned music," one of them showing a mechanical man tearing the strings out of a harp while an angel wept. They were not fools, and they were not Luddites. They were right. Those pit jobs did not come back. And yet (this is the part we have to be brave enough to admit), sound gave us the movie musical, the modern score, sfx, sound design, audio engineering, and an art form vastly larger than the one it disrupted. And it helped keep Hollywood in the forefront of world entertainment for the rest of the century and into the next. The loss was real. And yet the art form expanded. This is a story that has been told over and over again. To resist technology is to risk irrelevance. Just look at Kodak or Blockbuster. To embrace technology is to open doors of new possibility. Just consider Apple and Netflix. What I Learned From Walt Disney In the mid-1980s, I was tapped to lead Disney's animation division at a moment when the studio was at an inflection point. Animation wasn't just another business unit. It was the soul of the company, a medium revered because of Walt's genius and his passion. But the production system was cumbersome and unforgiving. A single movie was 125,000 individual hand-drawn and painted cels, photographed one frame at a time. Every revision carried a cost measured in months. These degrees of difficulty shaped the kinds of stories we could tell. We found our way forward in an unexpected place: Walt himself. The Disney archives held astonishing recordings of Walt explaining his creative process. His own writings. His notes and storyboards. Work product captured at every stage of his process. This was truly a gift. Listening, reading, sitting with the work itself, we heard him talk about character, about emotion, about how an audience feels when a character truly comes alive. He talked about making bold choices and refining a scene until it genuinely moved people. We didn't hear a word about pencils or paintbrushes. In fact, Walt was famous for being a technologist, forever hunting for state-of-the-art tools, often inventing them himself to achieve the images he saw in his head. But he never defined animation by the tools. He defined it by whether the audience believed the character. His principles were timeless. The tools were not. That realization changed everything. We co-developed the Computer Animation Production System (CAPS) with a young Northern California company called Pixar, replacing hand-painted cels with CGI. In The Little Mermaid, the final scene shimmered with a dimensionality and light that the old process simply couldn't achieve. In Beauty and the Beast, the ballroom sequence moved with a cinematic sweep that placed the audience inside the emotion of the moment. In Aladdin, the Cave of Wonders felt vast and alive, and the Magic Carpet became an intricate, compelling character all its own. In The Lion King, the stampede carried a scale and intensity that raised the emotional stakes beyond anything we'd done before. Technology didn't diminish the craft, it expanded the canvas. It gave artists more room to create. A decade later, the canvas expanded again. When Disney released Pixar's Toy Story, it wasn't simply a technical milestone. It was proof that a fully computer-animated film could carry real emotional weight, that it could make audiences laugh, cry, and believe. At DreamWorks, we made the difficult decision to sunset hand-drawn animation and become a fully computer-animated studio. It was the right thing to do, but it was not without pain. It cost talented people their place in an industry where they had worked their whole lives. Some made the leap to the new tools and did the finest work of their careers. Some never did. Tools are never the point. The instruments change with every generation. What endures is taste and imagination. The magical ability to make an audience feel. One of the greatest storytellers of our generation, George Lucas, succinctly captured the eternal essence of this issue: “It’s not the how, it’s the why.” A Distinction With a Difference I asked one of the leading AI models a question that has been challenging me for months. What is the difference between reasoning and creating? Its answer changed how I think about almost everything happening in this industry. It said . . . Reasoning and creating are two distinct cognitive modes, though they also work together. Reasoning is fundamentally evaluative and analytical. It operates on what already exists: facts, premises, evidence. It moves toward a conclusion that was in a sense already implied by the input. Reasoning is constrained by logic and truth. Its goal is to arrive somewhere correct, not to invent somewhere new. Creating is fundamentally generative. It produces something that didn't exist before. And crucially, there's no single right answer waiting to be found. A blank page has infinite valid responses. Creation involves choices that can't be fully justified by logic alone. Taste, intuition and vision fill the gap where deduction runs out. Reasoning is what Silicon Valley has been perfecting. Creating is what Hollywood has been practicing for more than a century. AI today operates almost entirely on the reasoning side of the line. It can deduce, evaluate, optimize, and pattern-match brilliantly. And while it can create, there is a real distinction to being creative. What it doesn’t yet have is those things that make us human: empathy, devotion, serendipity, the kind of creativity that comes from a person trying to say something only they could say. When the bot generates a piece of art, it is not trying to communicate anything. It is statistics, not soul; it is emulating things that have been done. By contrast, human creativity isn’t about repeating patterns of zeros and ones; it is about doing something new. One day, AI may close this gap. Three years ago, the leaders building AI would have called what they are achieving today, improbable, if not impossible. Impossible is no longer improbable. Today, the line between reasoning and creating is real. Even the leading technologists acknowledge we are not there yet. There is no scientific path to crossing this divide that anyone in the field can articulate today. Understanding that gap is where we will find common ground. A Path Forward In 2016, I closed one chapter in Hollywood with the sale of DreamWorks and opened another in Northern California, co-founding WndrCo. We’ve backed more than 50 founders building the next generation of technology and watched how breakthroughs in Silicon Valley emerge, first as experiments, then as platforms, and finally as infrastructure that reshapes entire industries. It's worth remembering that the last great revolution in animation also came from the north. Pixar was a Northern California company, forged not in the conventions of the Hollywood studio system, but in the technological breakthroughs of Silicon Valley. I've spent years on both sides of this bridge. For sure, I don’t have all the answers (take Quibi, for one!). But, from my past and present vantage points of my long career, here is what I see . . . Brilliant people in Northern California building this technology have made something extraordinary. They have earned the right for the rest of us to be, if not believers, at least optimistic that what comes next will be remarkable. But they have not made an artist. The tools are powerful, but they are not what makes a story matter. That knowledge lives 350 miles to the south, inside people whose life's work has informed the very models you are building. The right path forward includes them by design, with credit, with consent, and with compensation. Build this with the storytellers. Not on top of them. Taste is not something that can be synthesized, it is uniquely human. At the same time, Hollywood needs to accept that AI is not going away. The energy they are spending trying to make it disappear is energy they are not spending deciding the terms on which it will exist. And the terms are everything. The north needs something from it that they cannot build and cannot buy: creativity. The kind that takes a blank page and conjures a single right answer where there was none and has held audiences for a century. Without it, the most powerful reasoning engine ever invented will still be missing the only thing that makes a story worth telling. The artists who learn to wield these new instruments will do things the engineers never dreamed of. They always have. Edison invented the motion picture but made terrible movies. It took Chaplin, Lloyd, Keaton and so many others to make movies emotional. Now, the canvas is about to expand yet again. We should decide now that we intend to paint on it. There are so many valuable lessons in history. This has happened many times before, and it was never settled by the technology. It was settled by the terms. Sousa did not stop the phonograph; he helped write the law that made sure composers got paid. And two years ago, when the writers and the actors walked out, they were fighting for the very things Sousa was fighting for in 1906. Consent, compensation, the basic recognition that human creative work has a price that must be paid. The terms of that fight are still being negotiated, but the principle is older than any of us. The tools-versus-no-tools argument is a trap. First, we must all agree that there should be terms. Then we can have the crucial debate about what fairness requires. What I Learned From Steve Jobs Years ago, Steve Jobs said, "It's in Apple's DNA that technology alone is not enough. It's technology married with the liberal arts, married with the humanities, that yields us the result that makes our hearts sing." He was describing a device. But he could just as easily have been describing this tale of two cities. What I See Coming Soon As the barriers and the costs come down, more films will get made, not fewer. Studios will get to take more risks. There will be more seats at the table, and very soon entirely new forms of storytelling. In the 1980s, animation was dismissed as a niche corner of the business. Today it is one of the most beloved and profitable forms of storytelling in the world. In live action, filmmakers like Steven Spielberg, James Cameron and Peter Jackson embraced new visual tools not as shortcuts, but as instruments, and expanded cinema in the process. Every time storytelling has met a genuine technological shift, from synchronized sound to color to computer animation, it has redefined the boundaries of the medium and grown larger in the process. Assuredly, I don’t have all the answers, but I am confident that the creative opportunities will expand yet again. How we come through this is a choice. The north has the new tools. The south has the creative soul. The best future will draw on the best of both worlds.

  3. Google Blog:AI(RSS)68

    Google Vids 集成 Gemini Omni 1.1 Flash,所有账号可免费生成 1080p HD 视频

    Google 宣布所有 Google 或 Google Workspace 账号均可通过 Google Vids 使用最新的 Gemini Omni 1.1 Flash 模型免费生成高质量视频,入口为 vids.new 并选择 "Create AI videos"。

    推荐理由:原文来自产品经理宣布,交代了免费开放入口、具体模型和新控制功能,读者可据此判断是否改变自己的视频制作流程。

9月23日周三
  1. Josh Woodward39

    重大里程碑! @FlowbyGoogle: 每月有超过 2500 万人使用 Google Flow 来构思新点子、创作故事、打造酷炫作品。 谢谢大家。 我们将继续为所有用户提供每天额外 50 次额度。快来继续创作吧。

    引用Google Flow@FlowbyGoogle

    More than 25 million people are using Google Flow every month to dream up new ideas, create stories, and build cool things. Thank you. We’re continuing the 50 additional daily credits for all users. Dive in and keep creating.

9月22日周二
9月21日周一
  1. MiniMax Design (H3)36

    🔥社区从不停下折腾的脚步。 不只是基于 H3 做开发,还在不断深入内部,寻找让它更聪明的新方法。

    引用Kamimoto(かみもと)@sep_is_heim

    流行のJevをMiniMax H3に組み込んで、動画生成を高速化してみた!Attention処理のスパース化にJevを使用。 ・層ごとにJevが重要度を判定(4step 49層が対象) ・Jevがスパース率1%, 3%, 5%, 10%を選択 RTX4070で6分7秒→3分34秒で41.7%短縮!動画生成中にJevクラウドに問合せしているのに速い!

  2. elsewhere:文章(RSS)56

    葬AI评世界模型热潮:上不了大模型桌的人才另开一桌

    葬AI发文批评世界模型已成炒作概念,认为其起因是一批做不了大语言模型竞争的团队生造赛道,代表人物李飞飞和杨立昆做的其实是不同的东西。文章归纳了四类讲世界模型故事的人,称多数产品只是后训练开源视频模型,且因MiniMax H3开源可后训练才近期密集宣发,并预告将推出直播间Bench实测实时生成视频模型。

9月20日周日
9月19日周六
  1. Grok55

    Grok 官方推广其 Imagine 功能,称用户可以用 Grok Imagine 制作自己的电影场景。其引用的案例称 PJaccetturo 及其工作室在 9 天内制作了一个试播集,用了 3,627 张图像、3,043 个视频,生成成本 $2,677。

    引用Grok Imagine@imagine

    Hollywood would spend 7 figures on this Odyssey scene. @PJaccetturo and his studio made this pilot in 9 days, with 3,627 images, 3,043 videos, and generations costing $2,677. Here’s their breakdown of how they did it 🧵

9月18日周五
  1. MiniMax Design (H3)54

    MiniMax Design 转发用户演示:用 MiniMax H3 Max 的 r2v 功能将一张 3x3 分镜图一次性生成为视频,设置 480p、15 秒,提示词为「左上から右下のパネルにカットが切り替わる2Dアニメーション、複数パネル禁止、BGM禁止」(按左上到右下面板切换镜头的 2D 动画、禁止多面板、禁止 BGM),Quality 调整、参照强度标准,作者称生成画面与分镜设定基本一致。

    引用852話(hakoniwa)@8co28

    この3x3画像を1枚 r2v でAI動画化すると以下の設定でおおよそそのまま映像になる Minimax H3 Max r2v 480p 15秒 Prompt調整:Quality 参照強度:標準 「左上から右下のパネルにカットが切り替わる2Dアニメーション、 複数パネル禁止、BGM禁止」

  2. MiniMax (official)40

    Nunchux AI 与多校研究者推出 VC-Attention,为 MiniMax-H3 带来免训练低比特注意力加速,在 B200 上比 FlashAttention-4 快 1.6×、B300 上快 1.5×,保真度优于 SageAttention2。

    引用Nunchux AI@NunchuxAI

    Introducing VC-Attention: fast and accurate low-bit attention without retraining. On MiniMax-H3, VC-Attention speeds up attention by 1.6× on B200 and 1.5× on B300 over FlashAttention-4, with better fidelity than SageAttention2. It also works with existing sparse attention methods. Two key innovations: • V-Smooth reduces value quantization error. • ExpCast-FP8 speeds up softmax. Nunchux Attention, our proprietary extension, pushes the speedup to 1.9× on B200 and 1.8× on B300. Blog: http://www.nunchux.ai/blog/attention-is-the-video-bottleneck Technical Report: http://arxiv.org/pdf/2609.15810 Joint work by researchers at MIT, CMU, UC Berkeley, Stanford, and NVIDIA.

9月17日周四
9月16日周三
  1. Hao AI Lab42

    在 ComfyUI 上体验 FastVideo 的 FastH3 V2 🚀🚀

    引用ComfyUI@ComfyUI

    FastH3 by FastVideo is now available in ComfyUI Video and native stereo audio, generated together, in seconds. Best for: → Previz and animatics that need lots of takes, quickly → Timing and dialogue tests before committing to a full-quality render → Short-form and social work on tight turnarounds Run it locally in ComfyUI and soon on Cloud ⬇️