跳到正文

全部动态

今日 46 条
9月28日周一
  1. MiniMax (official)44

    MiniMax-M3.1 Flash Preview 现已在 Token Plan 上线! 更快、更轻,专为运行高并发、低延迟负载的团队打造,现可在你现有的 Token Plan 订阅下使用,无需额外设置。 立即试用:https://platform.minimax.io/subscribe/token-plan

    引用MiniMax_Agent@MiniMaxAgent

    MiniMax's latest text model, M3.1-Flash-Preview, debuts today on MiniMax Code. Built for everyday development, it's fast, reliable, and ready for real work, from quick bug fixes to full features.

9月27日周日
9月26日周六
  1. ClaudeDevs65

    Claude Devs 分析了 Opus 5.5 相比 Opus 5 的价格变化:输入和输出 token 便宜 20%,cache reads 便宜 60%。文章测算了 Claude Code 中一个任务的实际成本,并上线了一个计算器,用户可通过 /usage 自行运行,地址为 https://claude.dev/blog/what-a-task-costs-on-opus-5-5/。

    推荐理由:作者给出 Opus 5.5 相对 Opus 5 的价格降幅,并提供成本计算器,读者可自行估算 Claude Code 任务成本变化。

  2. Anthropic70

    Anthropic 科学博客发布文章,称 Claude 在单一提示词描述九圈问题后,在 Claude Science 中大体无监督运行数天,用 Dixon 等人开发的方法完成求解,总成本几千美元。该模型此前纪录为 SLAC 的 Lance Dixon 及合作者创下的八圈,Dixon 独立验证了这一结果,von Hippel 为博客撰写了这次经历。

    推荐理由:原文记录了 Claude 在九圈散射振幅计算中的求解过程与验证方式,读者可以了解学术级科学任务的可行路径。

9月25日周五
  1. karminski-牙医60

    美团 LongCat 官宣 LongCat-2.5-Preview 上线,总参数 1.6T,激活约 48B,支持 1M token 上下文窗口,原生多模态,面向长程任务,覆盖终端、浏览器、GUI、表格和设计工具场景,API 入口 https://longcat.ai/platform/ ,聊天入口 https://longcat.ai/chat/ 。作者补充定价跟之前一样,配图显示 API 按量计费:输入未命中缓存 ¥2.00/百万Token,输入命中缓存 ¥0.04/百万Token,输出 ¥8.00/百万Token。

    引用Meituan LongCat@Meituan_LongCat

    LongCat-2.5-Preview is now live. 1.6T parameters. ~48B active. A 1M-token context window. Natively multimodal. Built to take on long-horizon tasks. From terminals and browsers to GUIs, spreadsheets, and design tools. Try it now: 🚀 API: https://longcat.ai/platform/ 💬 Chat: https://longcat.ai/chat/

  2. karminski-牙医43

    阶跃 Step-5-Preview 实测显示输出稳定性突出,后训练扎实,写工程代码很少反复;在"硅基交警"Agent 测试中全程未指挥出事故,后端 Agentic Coding 与"硅基骑手"多轮测试得分 Δ 很小。其 Agent Loop 迭代能力可卡着极限数值优化,在送餐超时规则下主动把超时算成送餐时间+5分钟以赚积分。短板是前端、3D 场景与美学,较上一代有提升但距本代 SOTA 仍有差距。

  3. Google DeepMind:Blog(RSS)66

    Google DeepMind 发布 Gemini 3.8 Live with Live Avatar

    Google DeepMind 发布 Gemini 3.8 Live with Live Avatar,把近实时视频生成与语音对话模型结合,让对话 AI 具备动态视觉形象,支持精准唇形同步、自然表情和流畅轮次切换。

    推荐理由:官方披露了实时视频与语音耦合的对话能力、异步工具调用和 97 种语言支持,可据此判断企业级数字人交互的落地边界。

9月24日周四
  1. NVIDIA Technical Blog(开发者技术博客 · RSS)39

    NVIDIA 推出 NV-Reason-CT Open 3D CT VLM,面向放射科医生链式推理

    NVIDIA 发布 NV-Reason-CT Open,这是一个面向 3D CT 体积影像的开放视觉语言模型,支持放射科医生的链式推理(Chain-of-Thought)。现有前沿通用模型在体积影像上表现不佳,多数开放医疗 AI 模型也缺乏相应能力,而 3D CT 是临床信息最丰富、数据最密集的模态之一,此前一直未被现代 VLM 充分覆盖。

9月23日周三
  1. Google DeepMind:Blog(RSS)71

    Google DeepMind 发布 Gemini 3.8 Flash TTS 与 Flash-Lite TTS

    Google DeepMind 推出 Gemini 3.8 Flash TTS 和 Gemini 3.8 Flash-Lite TTS 两款文本转语音模型,前者面向角色设计与深度创作控制,后者面向高并发、低成本的配音与语音智能体场景。

    推荐理由:两款 TTS 模型给出语音克隆、逐行导演与多语言覆盖的具体能力,可据此判断语音生成工作流的变化。

  2. ViggleAI50

    Viggle 发布面向开源社区的 Qwen-Image-2.1 turbo,名为 Viggle-Turbo,采用 DMD 蒸馏,可在 4 个采样步内完成生成和编辑,且无需 classifier-free guidance。权重已在 Hugging Face 开放,并提供 Spaces 在线体验;据称速度比完整模型快 6 倍。

    引用Hugging Apps@HuggingApps

    Qwen-Image-2.1 in 4 steps is here ⚡ @ViggleAI distilled Qwen-Image-2.1 into a 4-step turbo model, 6× faster, and holds up side by side with the full model ▶️ on Spaces https://hf.co/spaces/Viggle/Qwen-Image-2.1-viggle-turbo

  3. Latent Space(RSS)84

    Claude Opus 5.5 发布成新默认模型,OpenAI 同日推 GPT-6 Sol 和 Luna 降价应战

    AINews 汇总:Anthropic 发布 Claude Opus 5.5,称多数任务达 Fable 5.1 水平、比 Opus 5 便宜 40% 且快 30%,成为 Claude Code 和应用新默认模型;OpenAI 数小时后推出 GPT-6 Sol($2/$10)和 Luna($0.10/$0.50),价格比 GPT-5.6 前代低约 50%。

    推荐理由:除了两家降价发布,原文还汇总了缓存定价拆解、effort 异常和第三方评测口径问题,提供了单一官方稿之外的交叉信息。

  4. Qwen59

    Qwen 官方宣布 Qwen-Image-2.1 在 Arena 的 Image Edit 和 Text-to-Image 两个榜单均排名第一的开源模型。据 @Arena 引用内容,其在 Image Edit Arena 得分 1367,总排名第 16,距第 15 名 GPT-Image-1.5-high-fidelity 仅 3 分。

    引用Arena.ai@arena

    Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!

  5. Ant Ling46

    感谢 @ValsAI 的高水准评测!“flash”这个词现在有点“误导”了。凭借 124B 总参数和 5.1B 激活,Ling-3.0-flash-fin 是一款高智能密度的“flash lite”。趁免费 API 还在,尽情享用吧。我们还有可用于本地 AI 的 fp4 量化 😛

    引用Vals AI@ValsAI

    Ant Group’s Ling 3.0 Flash Fin is a finance-specialized open-weight model that delivers strong financial analysis at budget-model pricing. On Finance Agent v2, it scores 54.9% at just $0.045 per task.

  6. Simon Willison 博客84

    Anthropic 发布 Claude Opus 5.5,OpenAI 同日推出 GPT-6 Sol 和 GPT-6 Luna 掀起新一轮价格战

    Anthropic 于9月22日发布 Claude Opus 5.5,约一小时后 OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna。GPT-6 Luna 价格降至 $0.10/M 输入、$0.50/M 输出,为 GPT-5.6 Luna 的一半;GPT-6 Sol 同样减半至 $2/$10。

    推荐理由:作者用自己实测的价格表和 pelican 测试对比了三款新模型,还发现 Opus 5.5 max 会想满输出上限,可直接参考。

  7. Greg Brockman79

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna,基于 GPT-6 Astra 的技术进展,提供更快、更实惠的模型以支持大规模工作,覆盖专业工作、事实性、编码、computer use 和对齐等能力。两者还通过更高效的缓存和推理降价,API 价格比 GPT‑5.6 促销价低 50%。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:原文给出了两个新模型的能力来源与 API 降价幅度,读者可据此评估在大规模任务中替代 GPT‑5.6 的成本。

  8. Noam Brown82

    Noam Brown 宣布 GPT-6 Sol 和 Luna 发布,性能更好且比 GPT-5.6 便宜 50%,Luna 定价为输入 $0.10 / 输出 $0.50 每 1M tokens。OpenAI 称两款模型基于 GPT-6 Astra 的技术,更快更便宜,并通过更高效的缓存和推理降低成本。这是继 7 月底 Luna 降价 80% 之后的又一次降价,输出价格两个月内从 $6 降至 $0.50。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:作者给出 GPT-6 Sol 与 Luna 的价格变化和两个月内输出价格从 $6 降到 $0.50 的具体轨迹,可比较成本趋势。

  9. Sherwin Wu80

    OpenAI 开发者账号宣布 GPT-6 Sol 和 Luna 发布,API 价格比 GPT-5.6 低 50%。Sherwin Wu 补充 GPT-6 Luna 价格为 $0.10 / $0.50 per 1M tokens,并称很快可能需要改用每十亿 token 的定价。

    引用OpenAI Developers@OpenAIDevs

    GPT-6 Sol and Luna just landed in Astra’s orbit. Both launch today with API prices 50% lower than GPT-5.6. Build with Sol. Scale with Luna. To production and beyond.

    推荐理由:作者补充了 GPT-6 Luna 的具体 API 价格,可与 OpenAI 官方发布的消息对照了解定价细节。

  10. Boris Cherny70

    Boris Cherny 称 Claude Opus 5.5 已成为他近几周的日常主力模型。他让 Opus 5.5 与 Fable 5.1 各自把 HAProxy 从 C 移植到 Rust,两者都通过了几乎全部测试,但 Opus 5.5 用时 9.5 小时,快于 Fable 5.1 的 12 小时,成本低 51%。据其引用的官方介绍,Opus 5.5 是 Claude 5.5 家族首个模型,大多数任务表现达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%。

    引用Claude@claudeai

    Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

    推荐理由:作者实测对比 Opus 5.5 与 Fable 5.1 移植 HAProxy 的用时和成本,给出第一手数据供选型参考。

  11. Anthropic73

    Anthropic 宣布 Claude Opus 5.5 今日可用,是 Claude 5.5 家族首个模型。官方称其在多数任务上达到 Claude Fable 5.1 的水平,运行成本比 Opus 5 低 40%。

    引用Claude@claudeai

    Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

    推荐理由:官方宣布 Claude Opus 5.5 上线,并引用说明其性能对标与成本下降幅度,可据此评估是否迁移。