跳到正文

#多模态

今日 2 条
9月10日周四
  1. Hugging Face:Blog(RSS)61

    Hugging Face 用 Gradio Workflow 重建 AUTOMATIC1111,推出 Workflow1111

    Hugging Face 发布 Workflow1111,用 gr.Workflow 在单个画布上以 73 个节点重建了 AUTOMATIC1111 的 11 条媒体管线,涵盖文生图、hi-resolution fix、图生图、VLM 反推提示词、检测生成 inpaint 蒙版、ControlNet 风格 annotator、背景去除、PNG Info 和图生视频。

    推荐理由:官方用 Gradio Workflow 在单个画布上复刻了 AUTOMATIC1111 的主要功能,读者可以对照它了解节点式工作流与 ComfyUI 的差异。

  2. Apple:Newsroom(RSS)61

    Apple 发布 Apple Watch Series 12 与 Ultra 4 健康功能,重新设计的 Health app 将引入 Longevity 标签页

    Apple 宣布为 Apple Watch 和 iPhone 推出升级的健康与健身体验。Apple Watch Series 12 和 Apple Watch Ultra 4 引入新 Health Sensing System。

    推荐理由:官方新闻稿列出新健康传感系统、Longevity 标签页和在家运动评估等功能的设备要求,可帮助读者判断升级与自家设备的关联。

9月4日周五
  1. Google Research62

    Google 与 HHMI Janelia 发布完整雄性果蝇脑连接组图谱

    Google Research 与 HHMI Janelia 及剑桥等机构合作,在 Cell 发表论文,发布完整雄性果蝇脑与中枢神经系统连接组图谱,包含超过 166,000 个神经元和 1.25 亿个突触连接,是迄今按神经元数量计最大的脑图谱。

    推荐理由:读者可了解 AI 重建如何把电子显微镜切片拼成完整脑图谱,以及这一资源对神经科学研究的用途。

9月3日周四
9月2日周三
  1. Google DeepMind:Blog(RSS)70

    Google DeepMind 为 Gemini 推出 agentic 视频理解功能

    Google DeepMind 推出 agentic video understanding,覆盖 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite,通过智能体循环动态调用原生视频工具按需检索画面、音频和字幕,而非固定帧率静态处理。

    推荐理由:原文给出了具体降本增效数字、适用模型和接入方式,开发者可据此评估是否切换视频分析流程。

9月1日周二
8月28日周五
8月27日周四
  1. Saining Xie43

    很高兴看到 RAE 扩展到视频!

    引用Minghui Guo@MinghuiGuo77

    🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE

8月26日周三
8月21日周五
8月20日周四
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)31

    用 NVIDIA FLARE 构建联邦多模态 AI 工作流

    NVIDIA 发布技术博客,介绍如何用 NVIDIA FLARE 构建联邦多模态 AI 工作流,在数据无法集中到一处的情况下跨数据本地站点协调视觉语言模型(VLM)训练。VLM 可支持视觉问答、图像描述和图文推理等任务,而联邦学习为这类分布式数据场景提供了训练协调方案。

8月17日周一
8月13日周四
  1. Microsoft Research 博客(RSS)37

    MindTopo 揭示多模态大模型的空间推理能力短板

    微软研究院推出 MindTopo 基准,从连续性、分离、顺序、包围、绳结五类拓扑关系评估多模态大模型的推理与规划能力。测试显示,模型在静态图像识别上表现明显优于交互式规划任务,失败多发生在规划阶段而非感知阶段,且整体远低于人类水平。图像与视频生成仅在单帧关系可见时偶有帮助,跨多步动作时难以维持拓扑约束。

8月12日周三
  1. Google DeepMind:Blog(RSS)68

    Google DeepMind 发布手语转文本模型 SL2T,落地 Gboard 与 Live Transcribe

    Google DeepMind 发布大规模多语言手语转文本模型 SL2T,为 Gboard 和 Pixel 11 的 Live Transcribe 带来手语转文字听写,首先支持 ASL 转英语,后续将扩展更多语言和设备。

    推荐理由:原文解释了手语翻译的两类核心难题和 SL2T 的技术取舍,并给出基准分数与隐私设计,读者可了解可用手语 AI 的实现路径。

8月11日周二
8月10日周一
8月4日周二
  1. Mistral AI:News(网页)62

    Mistral 发布 3B 开源多模态安全分类器 Shieldstral,Apache 2.0 开放权重

    Mistral 发布 Shieldstral,一个 3B 开源权重多模态安全分类器,将内容审核建模为策略自适应的问答任务,在推理时接受自然语言策略,无需重训即可适配文本与图像审核。它以 Apache 2.0 协议开放权重,在文本安全上匹敌最大 7 倍于其体积的开放审核模型,可在单张 16GB NVIDIA GPU 上运行,并输出经校准的连续安全分数。

    推荐理由:原文说明了将审核策略写成自然语言问题即可在推理时切换的机制,读者可据此评估它替代固定分类审核模型的可行性。

7月31日周五
7月30日周四
  1. Google DeepMind:Blog(RSS)69

    Google DeepMind 发布 Gemini Robotics ER 2,强化视频理解、任务编排与多机器人协作

    Google DeepMind 发布 Gemini Robotics ER 2,面向机器人的具身推理模型,支持视频理解、任务进度跟踪、工具编排与多机器人协作,并可将执行交给下层 VLA 模型。

    推荐理由:官方发布给出进度分类、moment finding 等具体数字和公开 API 入口,读者可以据此评估其机器人编排与视频理解能力。

7月28日周二
7月21日周二
7月16日周四