跳到正文

#多模态

今日 32 条
9月15日周二
  1. Apple:Newsroom(RSS)82

    Apple 发布 Siri AI 与新一代 Apple Intelligence 功能

    Apple 于 2026 年 9 月 14 日发布新一代 Apple Intelligence,推出全新版本的 Siri AI,具备个人上下文理解、屏幕感知、更广泛的系统级应用操作,并在 iPhone 相机、iPad 截图、Mac 快捷键和 Apple Vision Pro 中集成 Visual Intelligence。

    推荐理由:官方完整列出了 Siri AI 的具体能力、设备要求和区域限制,读者可以据此核对自己设备上的功能可用范围。

9月14日周一
  1. Google AI55

    Google Labs 推出个性化 AI 产品 Dreambeans,可从 Gmail、Calendar、Search、Gemini App 和 Google Photos 人脸分组等来源收集信息,生成定制化故事,如提醒朋友生日、生成插画故事和礼物建议。故事就绪后通过应用内每日通知推送,点击插画卡片可查看详情和直达链接。该体验为自愿开启,受隐私过滤保护,用户可通过点赞反馈帮助系统学习偏好。

9月11日周五
  1. Sierra:Blog(RSS)42

    Sierra 推出下一代多模态智能体:随对话形态变化的界面

    Sierra 发布多模态智能体,将语音、文本和可视化整合进同一段对话,并自动判断何时切换模式,用户无需重来或重复表述。该智能体一次构建即可部署到所有渠道,视觉组件同样通用;其 MCP UI 集成支持把产品卡片、对比表格、日历和表单直接嵌入对话,组件由企业自行设计和托管,更新后自动同步,无需重新部署或为各平台维护不同版本。

9月10日周四
  1. SiliconFlow70

    DeepSeek-V4.1-Flash 在 SiliconFlow 上线,提供 Day 0 支持。模型为 552B MoE,prefill 阶段约激活 8B、decode 阶段约激活 16B,支持原生视觉与 1M 上下文窗口,KV cache 占用相比 V4 Flash 约缩小 4 倍,采用 MIT 许可证,主打高吞吐生产级推理。

    推荐理由:上线方直接给出参数结构、上下文窗口、KV cache 对比和许可证信息,读者可据此评估实际部署选型。

  2. Hugging Face:Blog(RSS)61

    Hugging Face 用 Gradio Workflow 重建 AUTOMATIC1111,推出 Workflow1111

    Hugging Face 发布 Workflow1111,用 gr.Workflow 在单个画布上以 73 个节点重建了 AUTOMATIC1111 的 11 条媒体管线,涵盖文生图、hi-resolution fix、图生图、VLM 反推提示词、检测生成 inpaint 蒙版、ControlNet 风格 annotator、背景去除、PNG Info 和图生视频。

    推荐理由:官方用 Gradio Workflow 在单个画布上复刻了 AUTOMATIC1111 的主要功能,读者可以对照它了解节点式工作流与 ComfyUI 的差异。

  3. Apple:Newsroom(RSS)61

    Apple 发布 Apple Watch Series 12 与 Ultra 4 健康功能,重新设计的 Health app 将引入 Longevity 标签页

    Apple 宣布为 Apple Watch 和 iPhone 推出升级的健康与健身体验。Apple Watch Series 12 和 Apple Watch Ultra 4 引入新 Health Sensing System。

    推荐理由:官方新闻稿列出新健康传感系统、Longevity 标签页和在家运动评估等功能的设备要求,可帮助读者判断升级与自家设备的关联。

9月9日周三
9月5日周六
  1. Fei-Fei Li57

    World Labs 创始团队 Fei-Fei Li、Justin Johnson、Ben Mildenhall 与 a16z 的 Martin Casado 深入讨论新发布的空间智能世界模型 Atlas。Atlas 以新视角预测为核心,统一像素级生成与重建,将数字化采集一个空间的 3D 表示所需照片从 100 到 300 张降至 3 张。对话还涉及机器人瓶颈在数据而非芯片、新视角预测是 AI-complete 等话题,视频见 https://www.youtube.com/watch?v=qn1QDDBnTA0。

    引用a16z@a16z

    World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: https://www.youtube.com/watch?v=qn1QDDBnTA0 @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado

9月4日周五
  1. Google Research62

    Google 与 HHMI Janelia 发布完整雄性果蝇脑连接组图谱

    Google Research 与 HHMI Janelia 及剑桥等机构合作,在 Cell 发表论文,发布完整雄性果蝇脑与中枢神经系统连接组图谱,包含超过 166,000 个神经元和 1.25 亿个突触连接,是迄今按神经元数量计最大的脑图谱。

    推荐理由:读者可了解 AI 重建如何把电子显微镜切片拼成完整脑图谱,以及这一资源对神经科学研究的用途。

9月3日周四
  1. elsewhere:文章(RSS)43

    对卷卷的3小时访谈:从抖音到AI 3D、成为制造业OS的野心、基础模型不会吞噬一切

    数美万物创始人兼CEO任利锋(卷卷)在近3小时访谈中回顾了从0到1孵化抖音的经历,并介绍了公司最新发布的Hi3D 3.0 2048³模型。他认为基础模型不会吞噬一切,实体制造仍需能产出"生产级"3D资产的模型,难点在于拆件、连接结构、材料适配与交付。数美万物的目标是从Maker OS走向制造业OS,把普通人的创造欲送进现实世界的生产管线。

9月2日周三
  1. Fei-Fei Li57

    Fei-Fei Li 宣布 World Labs 团队发布多模态世界模型 Atlas,称其从零训练、可生成像素级精准相机控制帧。Atlas 还能从单张输入图像重建大场景、通过重排视频帧模拟时空、从一张或多张图像原生输出 3D 空间,并将多张带位姿图像合成一致的 3D 世界,应用方向覆盖 VFX 到机器人。

    引用World Labs@theworldlabs

    Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.

  2. Google DeepMind:Blog(RSS)70

    Google DeepMind 为 Gemini 推出 agentic 视频理解功能

    Google DeepMind 推出 agentic video understanding,覆盖 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite,通过智能体循环动态调用原生视频工具按需检索画面、音频和字幕,而非固定帧率静态处理。

    推荐理由:原文给出了具体降本增效数字、适用模型和接入方式,开发者可据此评估是否切换视频分析流程。

9月1日周二
8月28日周五
8月27日周四
  1. Saining Xie43

    很高兴看到 RAE 扩展到视频!

    引用Minghui Guo@MinghuiGuo77

    🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE

8月26日周三