跳到正文

#模型发布

今日 43 条
9月29日周二
  1. Thariq64

    Anthropic 发布 Claude Sonnet 5.5,是 Claude 5.5 家族的第二款模型,相比 Sonnet 5 明显升级,运行速度提升超过 30%,多数工作成本最高降低 30%。作者 Thariq 表示 Sonnet 与 Opus 5.5 让高阶智能更易获得,建议在构建工作流时尝试 Sonnet 5.5,以缓解 projects、claude tag 和 dynamic workflows 等抽象的 token 成本顾虑。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

  2. Charlie Holtz71

    Charlie Holtz 转发 Anthropic 官方推文,宣布 Claude Sonnet 5.5 发布,是 Claude 5.5 家族的第二款模型。官方称其相比 Sonnet 5 是明显升级,运行速度提升超过 30%,多数工作场景成本最多降低 30%。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

  3. ClaudeDevs79

    Anthropic 发布 Claude Sonnet 5.5,是 Claude 5.5 家族第二款模型,相比 Sonnet 5 更智能、更快 30% 以上,多数工作成本最多降低 30%。作者建议用于修复 bug、快速迭代功能等范围明确的日常任务,Claude Code 用量也因此更耐用。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

    推荐理由:原文给出 Sonnet 5.5 相对 Sonnet 5 的速度、成本和适用场景,读者可据此判断是否切换日常 Claude Code 任务。

  4. Boris Cherny66

    Anthropic 发布 Claude Sonnet 5.5,是 Claude 5.5 家族第二款模型。相比 Sonnet 5,它速度提升超 30%,多数工作成本最高降低 30%。作者 Boris Cherny 附上用 Sonnet 5.5 在 Claude Code 中修复 bug 的视频演示。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

    推荐理由:作者以 Claude Code 修复 bug 的视频演示新模型编码能力,可借此直观感受 Sonnet 5.5 相对上代的实际表现。

  5. Anthropic73

    Anthropic 官宣 Claude Sonnet 5.5 上线,是 Claude 5.5 家族的第二款模型。相比 Sonnet 5,速度提升超过 30%,多数工作的成本最高降低 30%。

    引用Claude@claudeai

    Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Sonnet 5, runs more than 30% faster, and costs up to 30% less for most work.

    推荐理由:Anthropic 官方发布 Claude Sonnet 5.5,原文给出速度提升与降价幅度,读者可据此比较是否升级迁移。

9月28日周一
  1. Hugging Face:Blog(RSS)71

    H 公司发布 Holo4 系列智能体模型,含 27B 稠密与 35B-A3B MoE 两个版本

    H 公司发布 Holo4 系列智能体模型,包含 27B 稠密版和 35B-A3B MoE 版,两者均已上线 H Models API,权重以 BF16、FP8、NVFP4 和 4-bit GGUF 格式开源在 Hugging Face。

    推荐理由:Holo4 同时给出两种尺寸、跨 GUI 与 MCP 的统一接口和公开轨迹,读者可据此比较开源智能体与闭源前沿的成本差距。

  2. MiniMax (official)44

    MiniMax-M3.1 Flash Preview 现已在 Token Plan 上线! 更快、更轻,专为运行高并发、低延迟负载的团队打造,现可在你现有的 Token Plan 订阅下使用,无需额外设置。 立即试用:https://platform.minimax.io/subscribe/token-plan

    引用MiniMax_Agent@MiniMaxAgent

    MiniMax's latest text model, M3.1-Flash-Preview, debuts today on MiniMax Code. Built for everyday development, it's fast, reliable, and ready for real work, from quick bug fixes to full features.

9月26日周六
9月25日周五
  1. karminski-牙医60

    美团 LongCat 官宣 LongCat-2.5-Preview 上线,总参数 1.6T,激活约 48B,支持 1M token 上下文窗口,原生多模态,面向长程任务,覆盖终端、浏览器、GUI、表格和设计工具场景,API 入口 https://longcat.ai/platform/ ,聊天入口 https://longcat.ai/chat/ 。作者补充定价跟之前一样,配图显示 API 按量计费:输入未命中缓存 ¥2.00/百万Token,输入命中缓存 ¥0.04/百万Token,输出 ¥8.00/百万Token。

    引用Meituan LongCat@Meituan_LongCat

    LongCat-2.5-Preview is now live. 1.6T parameters. ~48B active. A 1M-token context window. Natively multimodal. Built to take on long-horizon tasks. From terminals and browsers to GUIs, spreadsheets, and design tools. Try it now: 🚀 API: https://longcat.ai/platform/ 💬 Chat: https://longcat.ai/chat/

  2. Google DeepMind:Blog(RSS)66

    Google DeepMind 发布 Gemini 3.8 Live with Live Avatar

    Google DeepMind 发布 Gemini 3.8 Live with Live Avatar,把近实时视频生成与语音对话模型结合,让对话 AI 具备动态视觉形象,支持精准唇形同步、自然表情和流畅轮次切换。

    推荐理由:官方披露了实时视频与语音耦合的对话能力、异步工具调用和 97 种语言支持,可据此判断企业级数字人交互的落地边界。

9月24日周四
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)39

    NVIDIA 推出 NV-Reason-CT Open 3D CT VLM,面向放射科医生链式推理

    NVIDIA 发布 NV-Reason-CT Open,这是一个面向 3D CT 体积影像的开放视觉语言模型,支持放射科医生的链式推理(Chain-of-Thought)。现有前沿通用模型在体积影像上表现不佳,多数开放医疗 AI 模型也缺乏相应能力,而 3D CT 是临床信息最丰富、数据最密集的模态之一,此前一直未被现代 VLM 充分覆盖。

9月23日周三
  1. Google DeepMind:Blog(RSS)71

    Google DeepMind 发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS

    Google DeepMind 推出 Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 两款文本转语音模型,前者面向角色设计与深度创作控制,后者面向高并发、低成本的配音与语音智能体场景。

    推荐理由:两款 TTS 模型给出语音克隆、逐行导演与多语言覆盖的具体能力,可据此判断语音生成工作流的变化。

  2. ViggleAI50

    Viggle 发布面向开源社区的 Qwen-Image-2.1 turbo,名为 Viggle-Turbo,采用 DMD 蒸馏,可在 4 个采样步内完成生成和编辑,且无需 classifier-free guidance。权重已在 Hugging Face 开放,并提供 Spaces 在线体验;据称速度比完整模型快 6 倍。

    引用Hugging Apps@HuggingApps

    Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo

  3. Qwen59

    Qwen 官方宣布 Qwen-Image-2.1 在 Arena 的 Image Edit 和 Text-to-Image 两个榜单均排名第一的开源模型。据 @Arena 引用内容,其在 Image Edit Arena 得分 1367,总排名第 16,距第 15 名 GPT-Image-1.5-high-fidelity 仅 3 分。

    引用Arena.ai@arena

    Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!

  4. Ant Ling46

    感谢 @ValsAI 的高水准评测!“flash”这个词现在有点“误导”了。凭借 124B 总参数和 5.1B 激活,Ling-3.0-flash-fin 是一款高智能密度的“flash lite”。趁免费 API 还在,尽情享用吧。我们还有可用于本地 AI 的 fp4 量化 😛

    引用Vals AI@ValsAI

    Ant Group’s Ling 3.0 Flash Fin is a finance-specialized open-weight model that delivers strong financial analysis at budget-model pricing. On Finance Agent v2, it scores 54.9% at just $0.045 per task.

  5. Simon Willison 博客84

    Anthropic 发布 Claude Opus 5.5,OpenAI 同日推出 GPT-6 Sol 和 GPT-6 Luna 掀起新一轮价格战

    Anthropic 于9月22日发布 Claude Opus 5.5,约一小时后 OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna。GPT-6 Luna 价格降至 $0.10/M 输入、$0.50/M 输出,为 GPT-5.6 Luna 的一半;GPT-6 Sol 同样减半至 $2/$10。

    推荐理由:作者用自己实测的价格表和 pelican 测试对比了三款新模型,还发现 Opus 5.5 max 会想满输出上限,可直接参考。