很高兴看到 Google 用 Gemini 4 给人们带来惊喜。前沿领域有更多实验室,对消费者(竞争)和世界(减少权力集中)都是好事。期待看到它在真实场景中的表现。
#模型发布
#模型发布
今日 2 条
Nathan Lambert@natolambertAI 评分2020
Dongxi 东锡 NLP@dongxi_nlpAI 评分3131
swyx@swyxAI 评分2828引用Sholto Douglas@_sholtodouglasalso important news we fixed the writing
jietang@jietangAI 评分2626你确定吗?找到最优模型规模很棘手:数据量、激活参数量、环境数量,以及目标推理成本。模型性能还取决于许多其他因素,每个因素都带来各自的变数。
引用Charlie O'Neill@oneill_cFable is probably ~2-2.5T parameters, not 10T. Kimi K3 is 2.8T params, trained on maybe 20–30k Blackwell-equivalents. It lands within spitting distance of Fable 5 in terms of capabilities (5, not 5.1). Anthropic has far more compute than Moonshot, better rl environments, better architecture and better optimizers and all of that adds to capability per parameter. So if Fable is only slightly ahead of K3 with this in mind, it's almost certainly a smaller model. GPT-5.5 and 5.6 are smaller still (I'll say more on that later)
Nathan Lambert:Interconnects(RSS)精选AI 评分7171 Nathan Lambert 解析 GLM-5.3 与中国实验室如何跟上前沿
Z.ai 发布 GLM-5.3,目前仅在编码计划中提供,即将上线 API 并在两周后于 Hugging Face 开放权重,模型约 750B 参数,在多个 agentic coding 基准上超越 Kimi K3,部分超越 Claude Fable 5 或 GPT-5.6-Sol。
推荐理由:作者以第一手分析解释中国实验室如何保持前沿,给出发布节奏、RL 环境数据产业和模型定位等可迁移的判断框架。
Andrej Karpathy@karpathy精选AI 评分7676引用Claude@claudeaiFable 5 is state-of-the-art on nearly all tested benchmarks, with exceptional performance in software engineering, knowledge work, scientific research, and vision. The longer and more complex the task, the larger Fable 5’s lead over our other models.
推荐理由:作者补充了基准之外的第一手使用感受,指出长难题求解能力跃升和安全护栏偏敏感,可作选型参考。
Sam Altman:Blog(RSS)精选AI 评分8787 Sam Altman 谈 GPT-4o 发布:免费提供最强模型并强调新语音模式体验
Sam Altman 发文点评 OpenAI 当日发布的 GPT-4o,强调两点:一是把世界最强模型免费开放给 ChatGPT 用户,无广告,未来靠其他付费项目支撑,希望让数十亿人用上优秀 AI 服务;二是新语音(和视频)模式是他用过最好的计算机界面,达到接近人类水平的响应时间和表达力,感觉像电影里的 AI。他提到未来将加入可选的个性化、访问用户信息和代为执行操作等能力。
推荐理由:作者作为当事方说明 GPT-4o 免费开放的初衷和新语音模式的使用感受,读者可了解 OpenAI 对这两点的官方表态。