跳到正文

#编码

今日 54 条
9月26日周六
  1. OpenAI:官网动态(RSS · 排除企业/客户案例)40

    Proaction 借助 Codex 提升销售 60%、每月节省 75+ 小时

    Proaction 使用 Codex 后销售转化率提升 60%,每月节省 40–60 小时工程时间和 25–33 小时创始人时间。其联合创始人 Colin Knudsen 每月用 Codex 在 30–45 分钟内构建 4–6 个定制交互演示,并借助 Granola、Gmail、Slack、Linear、GitHub、HubSpot 等插件统一处理日常工作。

  2. Noah Zweben64

    Claude Code Remote Control 开始向 Pro 用户推送,先以 10% 灰度逐步放开,Team 和 Enterprise 版即将跟进。作者转引的推送说明给出使用步骤:更新到 claude v2.1.58 以上,可尝试重新登录获取新的灰度标记,然后运行 /remote-control;作者本人内容是发布 Opus 5.5 制作的黏土动画版演示视频。

    引用Noah Zweben@noahzweben

    Rolling out Claude Code Remote Control to Pro users - because they deserve to use the bathroom too . (Team and Enterprise coming soon). 🧻 Rolling out to 10% and ramping 1. Update to claude v2.1.58+ 2. Try log-out and log-in to get fresh flag values. 3. /remote-control

  3. Claude55

    Claude 官方发文称 Claude Opus 5.5 已发布数天,并分享了用户的一些探索发现。其中引用案例是用户 @RyanSael 让 Opus 5.5 通过构建一个可交互的镜头实验室来讲解相机对焦,模型用 1 小时 26 分钟一次生成,API 花费 $25.66,成品可通过 https://lens.lab.sael.net 访问,移动对焦环可看到玻璃元件移动清晰平面。

    引用Ryan Sael@RyanSael

    I asked Opus 5.5 to explain camera focus by building an interactive lens lab Here's what it came up with after 1 hour 26 minutes in one shot, $25.66 API cost https://lens.lab.sael.net Move the focus ring and you can see the glass elements shift the sharp plane through the scene

9月25日周五
  1. karminski-牙医43

    阶跃 Step-5-Preview 实测显示输出稳定性突出,后训练扎实,写工程代码很少反复;在"硅基交警"Agent 测试中全程未指挥出事故,后端 Agentic Coding 与"硅基骑手"多轮测试得分 Δ 很小。其 Agent Loop 迭代能力可卡着极限数值优化,在送餐超时规则下主动把超时算成送餐时间+5分钟以赚积分。短板是前端、3D 场景与美学,较上一代有提升但距本代 SOTA 仍有差距。

  2. François Chollet44

    我认为,无论你移动到哪个抽象层级,软件工程的"难度"本质上是不变的,因为人类认知会适应新工具,直到能够充分发挥自身能力。 工具只是可供性,不是让工作消失的魔法棒。 伟大的软件工程以前极其艰难。尽管工作流程已大不相同,如今它依然极其艰难。

    引用Simon Willison@simonw

    The more time I spend working with coding agents, the more convinced I am that they make software engineering even harder We can do amazing things with them, but unlocking their full potential requires extraordinary discipline and knowledge

  3. Claude Code:GitHub Releases(RSS)37

    Claude Code v2.1.282 发布

    Claude Code v2.1.282 新增 maxProseWidth 设置,可限制宽终端中 Claude 正文宽度而表格与代码块保持全宽,并新增启动提示及 /status、claude doctor 中列出被忽略或关闭遥测的变量。

  4. Pragmatic Engineer(RSS)50

    RoR 创始人 DHH 再掀“手写代码之死”争论

    Ruby on Rails 创始人 David Heinemeier Hansson 在 Rails World 主题演讲中宣布,至少在 37signals,专业工作手写代码的时代已经结束。这一表态再度引发“手写代码之死”的争论,文章还提及 Amazon 和 Meta 在招聘方面遇到的困难。

9月24日周四
  1. GitHub Blog24

    GitHub Copilot 应用如何渲染超大 pull request

    GitHub Copilot 应用重建了 pull request 视图,用一个含 2200 个文件、超 100 万行改动和 400 多条行内评论的开源 PR 做压力测试。其做法是把文档高度拆成确定性的代码几何与动态评论块两套几何:代码行高提前精确计算,评论高度则按需测量、修正幅度小且锚定在用户当前查看位置,从而避免滚动跳动。

9月23日周三
  1. Latent Space(RSS)84

    Claude Opus 5.5 发布成新默认模型,OpenAI 同日推 GPT-6 Sol 和 Luna 降价应战

    AINews 汇总:Anthropic 发布 Claude Opus 5.5,称多数任务达 Fable 5.1 水平、比 Opus 5 便宜 40% 且快 30%,成为 Claude Code 和应用新默认模型;OpenAI 数小时后推出 GPT-6 Sol($2/$10)和 Luna($0.10/$0.50),价格比 GPT-5.6 前代低约 50%。

    推荐理由:除了两家降价发布,原文还汇总了缓存定价拆解、effort 异常和第三方评测口径问题,提供了单一官方稿之外的交叉信息。

  2. Simon Willison 博客84

    Anthropic 发布 Claude Opus 5.5,OpenAI 同日推出 GPT-6 Sol 和 GPT-6 Luna 掀起新一轮价格战

    Anthropic 于9月22日发布 Claude Opus 5.5,约一小时后 OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna。GPT-6 Luna 价格降至 $0.10/M 输入、$0.50/M 输出,为 GPT-5.6 Luna 的一半;GPT-6 Sol 同样减半至 $2/$10。

    推荐理由:作者用自己实测的价格表和 pelican 测试对比了三款新模型,还发现 Opus 5.5 max 会想满输出上限,可直接参考。

  3. Greg Brockman79

    OpenAI 发布 GPT-6 Sol 和 GPT-6 Luna,基于 GPT-6 Astra 的技术进展,提供更快、更实惠的模型以支持大规模工作,覆盖专业工作、事实性、编码、computer use 和对齐等能力。两者还通过更高效的缓存和推理降价,API 价格比 GPT‑5.6 促销价低 50%。

    引用OpenAI@OpenAI

    Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.

    推荐理由:原文给出了两个新模型的能力来源与 API 降价幅度,读者可据此评估在大规模任务中替代 GPT‑5.6 的成本。

  4. Boris Cherny70

    Boris Cherny 称 Claude Opus 5.5 已成为他近几周的日常主力模型。他让 Opus 5.5 与 Fable 5.1 各自把 HAProxy 从 C 移植到 Rust,两者都通过了几乎全部测试,但 Opus 5.5 用时 9.5 小时,快于 Fable 5.1 的 12 小时,成本低 51%。据其引用的官方介绍,Opus 5.5 是 Claude 5.5 家族首个模型,大多数任务表现达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%。

    引用Claude@claudeai

    Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

    推荐理由:作者实测对比 Opus 5.5 与 Fable 5.1 移植 HAProxy 的用时和成本,给出第一手数据供选型参考。

  5. Claude Code:GitHub Releases(RSS)65

    Claude Code v2.1.280 发布:新增 Claude Opus 5.5 默认模型

    Claude Code 发布 v2.1.280,新增 Claude Opus 5.5(claude-opus-5-5)并设为默认 Opus 模型,支持 1M 上下文,价格 $4/$20 per Mtok、缓存读取 $0.20/Mtok;Pro 和 Team Standard 计划默认模型也从 Sonnet 改为 Opus。

    推荐理由:原文列出该版本新增 Claude Opus 5.5 默认模型、MCP 描述长度可配置等改动,读者可对照修复清单决定是否升级。

9月22日周二
  1. elsewhere:文章(RSS)68

    阶跃 Step 5 Preview 实测评测:数据可视化与金融分析亮眼,泛化和审美仍有短板

    阶跃发布 Step 5 Preview,总参数量 600B、激活参数 27B,有视觉输入,官方称在 Artificial Analysis 上涨 44 分、单任务成本仅为 Claude Opus 5 的 1/8。作者与友人实测发现其在数据可视化、金融分析上表现不错,但泛化性、领域知识和审美偏弱,思考过程过长导致长任务耗时且易中断,且长上下文下安全指令遵循可被绕过。

  2. Xiaomi MiMo40

    💗

    引用Design Arena@DesignArena

    BREAKING: MiMo-V2.6-Pro by @XiaomiMiMo lands at #8 overall (#3 open-weight) on Design Arena with an Elo of 1338. This is an impressive 54-point and 22-position increase from MiMo-V2.5-Pro. MiMo-V2.6-Pro also reaches #4 overall in Website (#2 open-weight) and #6 overall in Agentic Frontend Development (#2 open-weight), showing strong performance across both direct generation and agentic coding. This places @XiaomiMiMo's new model among leading models such as GPT-5.6 Sol by @OpenAI, Claude Opus 5 by @Anthropic, and Kimi K3 by @MoonshotAI. Congratulations to the @XiaomiMiMo team on returning to a top-10 placement on Design Arena!

  3. Xiaomi MiMo56

    小米 MiMo 发布 MiMo-V2.6-Pro,据 Arena 数据以 1628 分(AutoEval)位列 Code Arena: WebDev 总榜约第 10 名、MIT 许可开源权重模型中约第 3 名,较上一代 MiMo-V2.5-Pro 的 1475 分提升 153 分。该成绩为早期 AutoEval 分数,由基于人类偏好数据训练的 Reward Model 自动投票产生,随真实人类投票增加可能变化。

    引用Arena.ai@arena

    MiMo-V2.6-Pro just landed @XiaomiMiMo back in the top 10 on Code Arena: WebDev, debuting at ~#10 overall, and ~#3 among open-weights models with an MIT license. It scores 1628 pts (AutoEval), tying Claude Fable 5 (High) and just ahead of Hy4-preview (1624 pts). That's a +153 pt jump from the previous MiMo-V2.5-Pro (1475 → 1628 pts). Among open-weights models, it lands at ~#3. Impressively only 7 pts behind Qwen3.8 Flash Next (#2) and 46 pts behind Kimi K3 Max (#1). Note: this is an early AutoEval score, in which a Reward Model trained on Arena’s human preference data casts automatic votes in place of live votes. We’ll continue to see how scores converge as more live human votes come in. Congrats to the @XiaomiMiMo team on this release!

  4. Xiaomi MiMo66

    小米 MiMo 发布 MiMo-V2.6 Pro 与 Flash 两款全模态模型,通过规模化强化学习训练。官方称 Pro 在多数 agent 基准上与 Claude Opus 5 和 GPT-5.6 Sol 相当,Artificial Analysis Intelligence Index 得分 46,为开源模型中最高;能力覆盖编码、computer use、3D 推理和创作。

    推荐理由:原文给出 Pro 与 Claude Opus 5、GPT-5.6 Sol 的 agent 基准对比和开源范围,可据此评估其相对位置。

9月21日周一
  1. 小米 MiMo:GitHub 新仓库(模型发布)56

    小米 MiMo 开源 mimoagent:百行代码智能体在 SWE-bench Verified 得分超 74%

    小米 MiMo 在 GitHub 开源 mimoagent,一个仅约 100 行代码的 AI 智能体,可解决 GitHub issue 或在命令行中辅助用户。项目主打极简设计,无需庞大配置和大型 monorepo,并在 SWE-bench Verified 上取得超过 74% 的分数。仓库地址:https://github.com/XiaomiMiMo/mimoagent