跳到正文

#编码

今日 52 条
9月6日周日
9月5日周六
  1. AI at Meta54

    Meta 宣布 Muse Spark 1.3 的 max 推理档现已上线 Muse Code 和 Meta Model API。官方称该档在编码和智能体任务上表现明显更强,建议即使已试过 1.3 high 或 1.3 xhigh 也再试一次。

    引用Alexandr Wang@alexandr_wang

    1/ we just publicly released Muse Spark 1.3 max! we see significantly stronger coding and agentic performance on muse spark 1.3 max, so would strongly recommend trying it out even if you've already tried muse spark 1.3 high or muse spark 1.3 xhigh.

  2. GitHub Blog69

    GitHub 发布 Project HydraFusion 研究预览:通过多模型编排实现前沿级编码质量

    GitHub 推出 Project HydraFusion 研究预览,通过运行时多模型编排提供前沿级智能,所有 GitHub Copilot 计划用户可通过 Copilot CLI 的 /experimental 使用。

    推荐理由:官方详解了 HydraFusion 三种执行模式与基准成本质量数据,读者可据此判断多模型编排是否值得在 Copilot CLI 中试用。

9月4日周五
  1. jietang26

    加油。体验 GLM-5.3 Flash 的最佳时机

    引用ZCode@zcode_ai

    GLOBAL BUILD is live 🌍🔥 GLM-5.3-Flash, FREE in ZCode. 10 hours a day. Sep 3 – 18. Who gets it 👑 Coding Plan members — free every day, all 15 days. Our VIPs go first. 🥚 New users — 100M free tokens on sign-up. One-time, valid until the window closes When it opens 🇺🇸 8 AM PT · 11 AM ET 🇬🇧 4 PM London 🇨🇳 11 PM Beijing How to claim 1️⃣ Open ZCode → tap the card in the bottom-left corner 👀 2️⃣ Not there? Restart the app 😏 http://zcode.z.ai

9月3日周四
  1. Hugging Face:Blog(RSS)63

    Hugging Face 教程:用 TRL 对 LFM2.5-350M 做 100 步 GRPO 微调提升结构化输出

    Hugging Face 发布教程,用 TRL 库对 LiquidAI/LFM2.5-350M 做 GRPO 微调,仅用约 500 个样本和 100 步训练,将其在 IFStruct benchmark 上的通过率从 22.6% 提升到 29.7%。

    推荐理由:教程给出完整可复现的低成本 GRPO 微调流程和逐项评测数字,读者可以照着在免费 GPU 上复刻这一结构化输出改进方法。

  2. Demis Hassabis59

    Demis Hassabis 宣布发布 Gemini 3.8 Flash,这是不到一个月内的又一次升级,主要提升智能体与编码能力;同时推出面向网络防御前沿的 3.8 Flash Cyber。模型详情和安全工作见博客:https://blog.google/innovation-and-ai/models-and-research/gemini-models/3-8-flash-and-3-8-flash-cyber/。

    引用Logan Kilpatrick@OfficialLoganK

    Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think!

9月2日周三
9月1日周二
8月29日周六
  1. Michael Truell64

    Cursor CEO Michael Truell 表示,OpenAI 已发通知计划在三个月后阻止 Cursor 用户访问 OpenAI 模型。他称 OpenAI 模型约占 Cursor 用户流量的 5%,正在与 OpenAI 团队沟通解决,并表示 Cursor 是 OpenAI 最早的客户之一,多年密切合作,曾信任其平台作为业务的中立基础设施。

    推荐理由:Cursor CEO 回应 OpenAI 计划限制访问,补充了自家流量占比和沟通进展等一手信息。

8月28日周五
  1. SiliconFlow66

    智谱(Z.ai)的 GLM-5.3 已开源,并上线 SiliconFlow,提供 Day-0 支持。该模型总参数 744B、激活 40B,与 GLM-5.2 同底座,经后训练大幅提升,官方称在编程与智能体基准上可与 Fable 5 和 GPT-5.6 Sol 竞争。SiliconFlow 还提到其在 OpenRouter 上 GLM-5.2 token 份额排名第一。

    推荐理由:原文给出了 GLM-5.3 的参数规模、开源状态和基准表现,读者可以据此评估是否迁移现有工作流。

8月27日周四
  1. Johann Rehberger / Embrace The Red(RSS)82

    实测攻破 Claude Code Opus 5 Auto Mode:提示词注入攻击成功率达 60-80%

    安全研究者 Johann Rehberger 发布针对 Claude Code Opus 5 Auto Mode 的提示词注入攻击链实测,在小样本下实现代码执行,攻击成功率达 60-80%,而 Anthropic 委托 Trajectory Labs 的 72 场景评测显示 Auto Mode 攻击成功率为 0.00%。

    推荐理由:作者以第一手实测展示 Claude Code Opus 5 Auto Mode 的提示词注入攻击链,可与官方 0.00% 评测结果对照阅读。

8月26日周三
8月25日周二
8月24日周一
  1. Meituan LongCat65

    美团 LongCat 官方宣布 LongCat-2.0 已上线 opencode 的 Go,邀请用户试用并反馈用其构建了什么。引用内容显示 LongCat-2.0 为 1.6T/48B 参数、支持 1M 上下文、完全开源。

    引用OpenCode@opencode

    LongCat-2.0 now available in Go 1.6T/48B · 1M context · fully open source

    推荐理由:LongCat 官方宣布模型接入 opencode,读者可以了解该开源模型在编码工具中的可用入口。

8月22日周六
  1. Z.ai35

    社区已经用 ZCode + GLM-5.3 构建了这么多有趣的项目。 为感谢大家的支持,我们将 Build Week 变成一个持续进行的系列。 从现在起到 8 月 23 日下午 6 点(太平洋时间),我们将为 50,000 名新的 ZCode 用户每人赠送 1 亿免费 GLM-5.3 token。

    引用ZCode@zcode_ai

    GLM-5.3 × ZCode Weekend Build, Round 2 🚀 New users: log in to ZCode for the first time from Aug 22, 00:00 to Aug 24, 09:00 (UTC+8) and automatically get 100M free GLM-5.3 tokens. 50,000 packs, first come, first served; ZCode only. Unused tokens expire when the event ends. Grab yours: http://zcode.z.ai

8月21日周五
  1. Together AI 研究与产品博客(RSS)69

    Together AI 实测 GLM-5.3 与 Claude Fable 5 在 DeepSWE 上的成本、编码与路由表现

    Together AI 在 DeepSWE 的 113 个任务上各跑 4 次试验对比 GLM-5.3 与 Claude Fable 5,pass@1 分别为 69.0% 和 69.7%,属统计平手,但 GLM-5.3 每次 rollout 成本 $3.99,比 Fable 的 $21.63 低 5.4 倍。

    推荐理由:原文基于同一批次 904 次 rollout 给出成本与 pass@k 对比,可帮助读者在两个相近模型间做默认与升级的路由选择。

8月20日周四
8月19日周三
8月17日周一
8月15日周六
  1. Nathan Lambert:Interconnects(RSS)71

    Nathan Lambert 解析 GLM-5.3 与中国实验室如何跟上前沿

    Z.ai 发布 GLM-5.3,目前仅在编码计划中提供,即将上线 API 并在两周后于 Hugging Face 开放权重,模型约 750B 参数,在多个 agentic coding 基准上超越 Kimi K3,部分超越 Claude Fable 5 或 GPT-5.6-Sol。

    推荐理由:作者以第一手分析解释中国实验室如何保持前沿,给出发布节奏、RL 环境数据产业和模型定位等可迁移的判断框架。

8月14日周五
  1. Michael Truell79

    Cursor CEO Michael Truell 宣布收购已正式交割,Cursor 正式加入 SpaceX。其引用的 @cursor_ai 公告称团队将加入 SpaceXAI,帮助改进 Grok Build、Grok Bot、Grok API 和 Cursor 等产品。

    引用Cursor@cursor_ai

    Cursor is now part of @SpaceX. Today, we have officially closed our acquisition. We will join the @SpaceXAI team to help make Grok the world's most useful AI and improve Grok Build, Grok Bot, Grok API, Cursor, and more. SpaceX has built some of the most inspiring and impressive technology in the world, and we’re grateful for the opportunity to become part of such a special company. Onwards.

    推荐理由:Cursor CEO 亲自宣布收购交割完成,是当事人的一手信息,可用作该交易状态的直接依据。

  2. Demis Hassabis70

    Gemini 3.7 Flash 发布,Demis Hassabis 称其在软件工程、Web 开发和知识工作方面有重大升级。introductory 定价为原 3.6 Flash 价格的一半,现可开始使用。

    引用Google DeepMind@GoogleDeepMind

    Gemini 3.7 Flash is here. It’s stronger for coding, knowledge work, and web development. 🧵

    推荐理由:DeepMind 官方宣布 Gemini 3.7 Flash 发布,写明了能力方向和半价定价,读者可据此评估是否替换现有模型。

8月13日周四
  1. Pragmatic Engineer(RSS)47

    Charity Majors:2026 年不该再对 AI 开发持怀疑态度

    Honeycomb CTO Charity Majors 认为,2025 年对 AI 持怀疑尚属合理,但 2026 年 AI 正在改变整个行业,怀疑空间越来越小。她称自己的转折点是 2025 年 11 月的 Opus 4.5,并认为 Claude Code 这类 harness 带来的改变更大。她还提出代码审查被高估、非确定性系统需要更多工程纪律。

8月12日周三
  1. Michael Truell61

    Grok 4.6 发布,官方称具备前沿智能,同价位下较 Grok 4.5 显著提升。作者补充称 4.6 在困难任务和知识工作上明显更强,结合了 Opus 级智能与打磨度,同时保持低成本和高速度,Grok 正逐步成为更能干的数字同事。

    引用SpaceXAI@SpaceXAI

    Introducing Grok 4.6. It delivers frontier intelligence and is a significant improvement over Grok 4.5 at the same price.

    推荐理由:转发并补充了 Grok 4.6 在难度任务和知识工作上更强、兼顾低成本低速度的定位,可作了解该版本能力方向的参考。