跳到正文

#多模态

今日 32 条
9月30日周三
  1. Rohan Paul60

    阶跃星辰(StepFun)与 ACE Studio 发布 StepAudio 3 Music,是其首个音乐生成基础模型,可根据提示词与歌词生成完整歌曲,支持歌曲生成、纯器乐生成、音乐翻唱和人声编曲四种工作流,基于 ABC-COT 先规划音乐结构与编曲再合成,交互 demo 见 https://static.stepfun.com/blog/stepaudio3/music/。

    引用StepFun@StepFun_ai

    StepFun and ACE Studio present StepAudio 3 Music, StepFun’s first music generation foundation model — built to turn a prompt and lyrics into a complete song. Describe the sound you want: • Genre, mood and vocal character • Instruments, key and BPM • Song structure and arrangement Four workflows in one model: 🎤 Song generation 🎹 Instrumental generation 🔁 Music cover 🎙️ Vocal-to-song arrangement Powered by ABC-COT, it plans musical structure and arrangement before synthesis. Generate a version, rewrite the prompt, edit the ABC notation, and iterate toward the sound in your head. (ABC-COT API coming soon) Try the interactive demo: › https://static.stepfun.com/blog/stepaudio3/music/

  2. NVIDIA Technical Blog(开发者技术博客 · RSS)30

    NVIDIA VSS Blueprint 3.3 发布:降低视觉 AI 智能体构建与运行成本

    NVIDIA 发布 Metropolis Blueprint for Video Search and Summarization(VSS)3.3,用于降低视觉 AI 智能体的构建与运行成本。该蓝图及其智能体技能将视频接入、流处理、事件检测、检索、摘要与报告整合为可维护的系统,让视觉语言模型在生产规模下理解视频。

  3. Apple:Newsroom(RSS)34

    Apple Creator Studio 将迎来新更新

    Apple Creator Studio 将推出全新模板、端到端电影级视频功能及多项提效特性。Freeform 新增自动适配深色模式、文件夹分组协作,Writing Tools 借助 Apple Intelligence 可改写、校对和摘要手写笔记,并支持通过 Shortcuts 自动添加内容。

9月29日周二
  1. Microsoft Research 博客(RSS)66

    微软研究院推出面向生物学复杂性的 AI 研究系统 Quine

    微软研究院推出 Quine,一个面向生物学复杂性的 AI 研究系统,由生物学世界模型和连接模型、科学工具、文献与研究者的交互式 harness 两部分组成。该系统与 Broad Institute 合作用于胰腺导管腺癌研究,预测并排序数千种化合物以推动肿瘤细胞状态转变,最高排名化合物在湿实验中产生了最大的预期转变,从缩小化合物搜索空间到选出候选验证仅用一个周末。

    推荐理由:微软研究院公开 Quine 的多模态世界模型与实验闭环设计,读者可了解 AI 参与生物实验设计与验证的具体路径。

  2. OpenBMB34

    MiniCPM-o 4.5 现已进入 SGLang Omni v0.1.7。 为开发者提供更多运行和构建该模型的灵活性。

    引用Guitar Cat + LLM@GenAI_is_real

    Hi everyone, today we released SGLang Omni v0.1.7. This release includes 75 merged PRs and welcomes 8 new contributors, with 8 first-time contributions. We added MiniCPM-o 4.5, NVIDIA PersonaPlex-7B, and OmniTyper powered by MLX streaming ASR, while further improving realtime and stateful Omni serving. 1.Performance: continued optimizations for Qwen3-TTS, Qwen3-Omni, CosyVoice3, MOSS-TTS, and AuK, covering Prefill CUDA Graph, speaker/reference encoding, kernel fusion, batching, and vocoder hot paths. 2.Serving: added Omni session lifecycle, the SGLang streaming session bridge, and a shared /v1/realtime WebSocket runtime, while further improving realtime ASR and streaming serving. 3.Models & hardware: added MiniCPM-o 4.5 multimodal input and speech output, plus PersonaPlex-7B offline speech-to-speech. MiniCPM-o and MiniMax-Music3 now support Intel XPU, with further MUSA support for Qwen3-TTS. 4.Runtime: improved breakable Prefill CUDA Graph, Talker / Code2Wav colocation, priority CUDA streams, scheduler admission, and profiling infrastructure to reduce host overhead and improve high-concurrency stability. https://github.com/sgl-project/sglang-omni/releases/tag/v0.1.7 https://github.com/sgl-project/sglang-omni

  3. The Verge:AI(RSS)80

    AMD 以约 82 亿美元全股票收购 World Labs

    AMD 宣布以约 82 亿美元全股票交易收购由李飞飞联合创立的 AI 研究实验室 World Labs,交易预计今年年底完成。World Labs 于 2024 年成立,数月内估值达 10 亿美元,2025 年推出首个商业产品世界生成模型 Marble,可根据提示词生成可交互 3D 世界。

    推荐理由:AMD 以约 82 亿美元全股票收购 World Labs,读者可了解交易结构、团队去向及其对 AI 硬件与模型协同的影响。

9月28日周一
  1. NVIDIA AI48

    工程师之间的合作最棒了 🙌 恭喜 @AIatMeta 推出 Muse Realtime Avatar!我们与他们的团队合作,让模型更高效,这样虚拟形象在你与它对话时就能跟上节奏。

    引用AI at Meta@AIatMeta

    @finkd just unveiled Muse Realtime Avatar, our real-time embodiment technology that turns Muse Realtime Voice into expressive, interactive avatars. Muse Realtime Avatar enables an entirely new range of interactions in @Muse, starting with real-time conversations. 🧵👇 #MetaConnect

9月27日周日
9月26日周六
9月25日周五
  1. karminski-牙医60

    美团 LongCat 官宣 LongCat-2.5-Preview 上线,总参数 1.6T,激活约 48B,支持 1M token 上下文窗口,原生多模态,面向长程任务,覆盖终端、浏览器、GUI、表格和设计工具场景,API 入口 https://longcat.ai/platform/ ,聊天入口 https://longcat.ai/chat/ 。作者补充定价跟之前一样,配图显示 API 按量计费:输入未命中缓存 ¥2.00/百万Token,输入命中缓存 ¥0.04/百万Token,输出 ¥8.00/百万Token。

    引用Meituan LongCat@Meituan_LongCat

    LongCat-2.5-Preview is now live. 1.6T parameters. ~48B active. A 1M-token context window. Natively multimodal. Built to take on long-horizon tasks. From terminals and browsers to GUIs, spreadsheets, and design tools. Try it now: 🚀 API: https://longcat.ai/platform/ 💬 Chat: https://longcat.ai/chat/

  2. Tripo30

    Tripo x CMU 游戏创作社团! 我们很高兴启动与 @CarnegieMellon 学生为期一学期的合作,探索 AI 原生游戏开发。 我们以一场工作坊开场,测试了一条完整流程: 创意 → GPT-6 Astra + Tripo + Blender + Codex → 可玩的 @unity 游戏。 在 2026 年秋季学期中,学生们将把 Tripo 整合进自己的项目。我们会全程提供技术反馈,最终在 12 月迎来最终展示和 Tripo Award。

  3. Google Research67

    Google Research 发布 AI 视频联合导演框架,实现连贯长视频生成

    Google Research 发布 AI video co-director,一个构建在 Gemini 和 Veo 之上的多智能体编排框架,用于生成连贯的多镜头长视频,并原生继承 SynthID 水印等安全机制。

    推荐理由:Google 把长视频生成拆成四个可组合的智能体框架,并给出三套自建基准与量化结果,便于对照现有方案。

  4. Google Blog:AI(RSS)43

    Chrome 升级学习习惯的 5 种方式:Gemini 支持播客视频解析与互动测验

    Chrome 将 Gemini 的媒体理解能力在桌面端扩展到播客及 YouTube 以外的视频,播放后即可提取要点、定位信息或解释难点。Gemini 还能基于标签页内容生成互动测验,已在美印桌面端以英文推出,未来数月扩展至更多地区和移动端。此外 Chrome 新增跨设备发送标签页时保留滚动位置和未填完表单,并支持分屏视图、标签组、沉浸式阅读模式和朗读功能。

  5. Google DeepMind:Blog(RSS)66

    Google DeepMind 发布 Gemini 3.8 Live with Live Avatar

    Google DeepMind 发布 Gemini 3.8 Live with Live Avatar,把近实时视频生成与语音对话模型结合,让对话 AI 具备动态视觉形象,支持精准唇形同步、自然表情和流畅轮次切换。

    推荐理由:官方披露了实时视频与语音耦合的对话能力、异步工具调用和 97 种语言支持,可据此判断企业级数字人交互的落地边界。

9月24日周四