Apple 发布 Siri AI 与新一代 Apple Intelligence,随 iOS 27 等系统更新推出
Apple 发布新一代 Apple Intelligence 和全新 Siri AI,后者以英文 Beta 陆续推出,10 月将支持法语、日语、韩语、葡萄牙语和西班牙语。
推荐理由:官方完整说明了 Siri AI 的能力范围、新 Apple Intelligence 功能和各区可用性差异,读者可据此判断升级价值。
Apple 发布新一代 Apple Intelligence 和全新 Siri AI,后者以英文 Beta 陆续推出,10 月将支持法语、日语、韩语、葡萄牙语和西班牙语。
推荐理由:官方完整说明了 Siri AI 的能力范围、新 Apple Intelligence 功能和各区可用性差异,读者可据此判断升级价值。
Apple 于 2026 年 9 月 14 日发布新一代 Apple Intelligence,推出全新版本的 Siri AI,具备个人上下文理解、屏幕感知、更广泛的系统级应用操作,并在 iPhone 相机、iPad 截图、Mac 快捷键和 Apple Vision Pro 中集成 Visual Intelligence。
推荐理由:官方完整列出了 Siri AI 的具体能力、设备要求和区域限制,读者可以据此核对自己设备上的功能可用范围。
Ling-3.0-flash-VL 两项更新: - 现已在 OpenRouter 上线,提供两周免费访问 - FP4 和 INT4 量化版本现已开源 通过 API 试用或自行部署——由你选择。
Sierra 发布多模态智能体,将语音、文本和可视化整合进同一段对话,并自动判断何时切换模式,用户无需重来或重复表述。该智能体一次构建即可部署到所有渠道,视觉组件同样通用;其 MCP UI 集成支持把产品卡片、对比表格、日历和表单直接嵌入对话,组件由企业自行设计和托管,更新后自动同步,无需重新部署或为各平台维护不同版本。
Gemini 应用现已登陆 Windows。 只需一个快捷键,助你保持专注。按下 Alt + Space,即可在你常用的工具和日常应用旁润色草稿、总结长文档、头脑风暴新想法,以及创建自定义图像和视频。
推荐理由:上线方直接给出参数结构、上下文窗口、KV cache 对比和许可证信息,读者可据此评估实际部署选型。
Hugging Face 发布 Workflow1111,用 gr.Workflow 在单个画布上以 73 个节点重建了 AUTOMATIC1111 的 11 条媒体管线,涵盖文生图、hi-resolution fix、图生图、VLM 反推提示词、检测生成 inpaint 蒙版、ControlNet 风格 annotator、背景去除、PNG Info 和图生视频。
推荐理由:官方用 Gradio Workflow 在单个画布上复刻了 AUTOMATIC1111 的主要功能,读者可以对照它了解节点式工作流与 ComfyUI 的差异。
MiniMax M3 在 AMD Instinct MI355X 上的 vLLM 服务性能大幅提升:SemiAnalysis InferenceX 基准显示,并发 32 时 MXFP8 标准服务从 109.1 升至 342.4 output tokens/s/GPU(3.14×),中位 TTFT 从 1.46 秒降至 0.67 秒。
NVIDIA 技术博客解析了 Encode-Prefill-Decode(EPD)分离架构这一多模态模型推理优化技术,它将视觉编码器阶段与 prefill、decode 阶段分离。该方案对图像密集提示词、中短输出和量化 MoE 模型最有效,配合 NVIDIA Dynamo 可实现最高 5 倍加速。
Apple 宣布为 Apple Watch 和 iPhone 推出升级的健康与健身体验。Apple Watch Series 12 和 Apple Watch Ultra 4 引入新 Health Sensing System。
推荐理由:官方新闻稿列出新健康传感系统、Longevity 标签页和在家运动评估等功能的设备要求,可帮助读者判断升级与自家设备的关联。
现在你可以在 Hugging Face 上体验 Ling 的 VL 模型了,由我们的 day0 合作伙伴 @novita_labs 提供支持
🤗 Novita now supports Ling-3.0-flash-VL on @huggingface. 🎁 Free for 14 days. • 124B total parameters · 5.5B active parameters per token • Native image and video understanding • Built for multimodal reasoning and agentic workflows
来 DeepInfra 试试最新的 Ling-3.0-flash-VL,我们的 day0 合作伙伴 @DeepInfra
Ling-3.0-flash-VL is live on DeepInfra — day 0 with @AntLingAGI. https://deepinfra.com/inclusionAI/Ling-3.0-flash-VL
One more thing…
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: https://www.youtube.com/watch?v=qn1QDDBnTA0 @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
我们的短视频概览国际扩展已在网页端向所有用户100%全量推送(移动端即将推出!) 70+ 种新语言和 3 种新英文变体,满足你所有的短视频需求。 你觉得怎么样?(我们自己挺满意的。)
Google Research 与 HHMI Janelia 及剑桥等机构合作,在 Cell 发表论文,发布完整雄性果蝇脑与中枢神经系统连接组图谱,包含超过 166,000 个神经元和 1.25 亿个突触连接,是迄今按神经元数量计最大的脑图谱。
推荐理由:读者可了解 AI 重建如何把电子显微镜切片拼成完整脑图谱,以及这一资源对神经科学研究的用途。
WeatherNext 3 是我们在全球天气预报方式上的一项重大突破。⛅ 该模型与 @GoogleResearch 联合开发,直接从真实世界的实时观测中学习,从而更快地给出更本地化、高度准确的预测。🧵
H Company 发布 NeoMME 系列多语言多模态编码器,提供 260M 和 800M 两种规格,用单个双向 Transformer 处理文本和 32×32 图像 patch,以 masked discrete-diffusion 目标从零预训练,不依赖预训练视觉塔或因果语言模型。
数美万物创始人兼CEO任利锋(卷卷)在近3小时访谈中回顾了从0到1孵化抖音的经历,并介绍了公司最新发布的Hi3D 3.0 2048³模型。他认为基础模型不会吞噬一切,实体制造仍需能产出"生产级"3D资产的模型,难点在于拆件、连接结构、材料适配与交付。数美万物的目标是从Maker OS走向制造业OS,把普通人的创造欲送进现实世界的生产管线。
Google 提出 MAPL-EMIT 深度学习框架,基于 EMIT 高光谱辐射数据自动检测、预测增强并定位全球甲烷羽流,在专家标注羽流上召回率达 84%。该模型采用 Swin-S 视觉 Transformer,同时完成增强量化、羽流分割与源定位三项任务,训练数据为注入真实 EMIT 场景的 360 万个合成甲烷羽流。
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
推出 Atlas: 全球首个多模态世界模型,可生成图像和视频帧,具备像素级精准的相机控制,并在 3D 中重建它们。 建模世界,移动相机,模拟空间与时间。
Google DeepMind 推出 agentic video understanding,覆盖 Gemini 3.7 Flash、3.6 Flash 和 3.5 Flash-Lite,通过智能体循环动态调用原生视频工具按需检索画面、音频和字幕,而非固定帧率静态处理。
推荐理由:原文给出了具体降本增效数字、适用模型和接入方式,开发者可据此评估是否切换视频分析流程。
上海人工智能实验室 InternLM 团队发布 InternLumina-U2,一个面向全视觉理解、图像生成与编辑的多码本扩散大语言模型。该模型将扩散生成能力与语言建模统一在同一框架内,可同时处理视觉理解与图像生成、编辑任务。
微软研究院联合华盛顿大学和 Providence 发布 GigaPath-Flash 与 GigaTIME-Flash,两者均以 Apache 2.0 许可在 Hugging Face 开放权重。
Hugging Face 与 Voice Arena 合作,在 Open ASR Leaderboard 引入 Monsoon en-IN 和 Monsoon hi-IN 两套评测集,其中印地语是该多语言榜单首个印度语言。
Google 在 Earth AI 之下推出实验性研究能力行星预测引擎(PPE),从自然语言查询出发自主完成数据发现、特征工程、模型训练与评估的全流程。
Google DeepMind 发布 Gemini Omni 1.1 Flash,通过 Gemini API 和 Google AI Studio 面向开发者提供新的创意控制与生成式视频能力。
推荐理由:原文来自官方,列出了各能力具体参数、速度和价格差异,开发者可据此评估是否接入自己的视频工作流。
🚀 What if video generators could build on representations that already understand the visual world? We are excited to introduce V-RAE: Rethinking Video Latent Spaces for Generation. Recent progress in image generation has begun to move beyond conventional VAE latents, exploring both direct pixel-space and representation-based approaches. Video generation, however, still depends heavily on latent compression, as the scale and redundancy of spatiotemporal data make direct modeling prohibitively expensive. However, most video VAEs are optimized for pixel reconstruction, and a latent space that reconstructs well is not necessarily easy to generate. V-RAE takes a different approach: it directly uses representations from frozen vision foundation models as the generative latent space, rather than as auxiliary supervision. We study DINOv3, SigLIP2, EUPE, and V-JEPA 2.1. A lightweight temporal attention pooling module compresses their dense features by 4×, followed by a spatiotemporal Transformer decoder. Under matched generation backbones, latent budgets, and training settings, V-RAE achieves: 🏆 2.13 rFVD on Kinetics-600 🎬 117.86 gFVD on UCF101 and 19.16 gFVD on Kinetics-600 ⚡ Up to 6× faster convergence than VAE-based latent spaces 🧠 90.92% semantic probing accuracy on UCF101 🌍 Better future prediction on Cityscapes, reducing gFVD from 144.47 to 111.36 Our experiments also reveal a broader finding: Good Reconstruction ≠ Good Generation. During generation, predicted latents inevitably deviate from real encoding trajectories. If the latent space is not sufficiently smooth, small errors can be amplified into visible artifacts. We therefore introduce tFVD to evaluate temporal smoothness and robustness to latent prediction errors. It correlates much more strongly with downstream generation quality, reaching 0.919 on Kinetics-600. The takeaway: A latent space is not merely where videos are compressed—it determines what the generator must learn. When semantics and temporal structure are already organized in the representation, generation becomes easier to learn. Representation first. Generation follows. Many thanks to my mentors, @ScottNLP and @SQWu_Tori, for their continuous guidance and support. I am also deeply grateful to @sainingxie for his valuable guidance and invaluable feedback, which greatly helped shape V-RAE. 🙏 Hi @_akhaliq, we would truly appreciate your help in sharing V-RAE with the broader AI research community. Thank you! 🙏 📄 Paper: https://arxiv.org/abs/2608.13556 💻 Code: https://github.com/V-RAE/V-RAE 🤗 Models: https://huggingface.co/Guomh0707/V-RAE-Models 🌐 Project: https://v-rae.github.io #VideoGeneration #GenerativeAI #ComputerVision #WorldModels #RepresentationLearning #RAE
Google 发布 GlucoFM,一个采用双流设计的自监督基础模型,将血糖的缓慢基线趋势与短期波动分离,并保留时间与缺失信息。模型在 109,066 小时无标注 CGM 数据上预训练,覆盖 477 条受试者/记录,在 4 个队列、7 项临床任务共 14 组评估中平均 PR-AUC 比最强 GluFormer 变体高 5.8 个百分点。
智谱(Z.ai)发布 GLM-5.3-Flash,为原生多模态模型,拥有 1M-token 上下文窗口,参数规模 320B-A18B,以 MIT 许可开源权重。
推荐理由:官方公布了定价定位、1M 上下文、MIT 许可和国产芯片适配等关键细节,可帮助读者评估这款轻量模型的实际可用性。