一位工程师爆料:大公司团队全靠 Claude Code 写代码,无人阅读
一位新入职大公司的工程师称,团队里的规格、代码、测试、PRD、工单及其处理、报告等全部由 Claude Code 生成,从 L1 到 L7 的工程师都在做同样的事——和 Claude 对话。没人喜欢这种方式,但被要求尽可能多地产出,管理层多次表示推代码不是瓶颈,员工每天工作 12 到 13 小时只为按回车,没有人阅读任何内容。
一位新入职大公司的工程师称,团队里的规格、代码、测试、PRD、工单及其处理、报告等全部由 Claude Code 生成,从 L1 到 L7 的工程师都在做同样的事——和 Claude 对话。没人喜欢这种方式,但被要求尽可能多地产出,管理层多次表示推代码不是瓶颈,员工每天工作 12 到 13 小时只为按回车,没有人阅读任何内容。
Simon Willison 发布 llm-keys-ui 0.1,一个用于在不把 API key 粘贴进智能体会话的情况下配置密钥的插件。
Claude Code v2.1.278 将 Claude API、Enterprise 用户以及 Bedrock、Vertex、Foundry 和网关上的 auto 模式默认切换为服务端分类器,不再收取分类器开销,计费回退时会给出警告。
Claude Code v2.1.277 新增 AGENTS.md 支持:项目无 CLAUDE.md 时会改读 AGENTS.md,可在 /config 的 "Project instructions" 中修改(Bedrock、Vertex、Foundry 暂不支持)。
GitHub Podcast 最新一期拆解了五个 AI 热门观点:AI 生成的代码仍需阅读和负责,只是审查力度应按风险分级;Skills 与 MCP 解决不同问题,MCP 提供工具与数据的标准接入,Skills 封装团队流程与最佳实践,二者可组合使用;RAG 并未消亡,检索能为模型提供训练数据之外的信息,减少 token 消耗并让回答更有依据。
a16z 图表周报引用近期论文和 SensorTower 数据指出,AI 代码生成工具让每月新应用数量在 iOS、Android 和 Chrome 上翻倍甚至翻两番,但下载量和评分基本停滞,达到 10+ 评分或 100+ 下载等规模的应用占比大幅下降。
MiniMax Code CLI 现已开放: https://github.com/MiniMax-AI/minimax-code
https://x.com/i/article/2100908964779040768
作者参加超级发电站主办的AI艺术黑客松,做了一个改编自麦浚龙与谢安琪概念专辑《The Album》的实时互动游戏,玩家扮演酒保自由输入回应客人。技术方案采用2.5D等距视角,由剧本、独立subagent演员、导演程序和预制美术资产四层组成,demo部署在 https://elsewhere.news/the-album 并凭邀请码有限开放。
Claude Code 发布 v2.1.276,修复了当 ANTHROPIC_BASE_URL 指向代理或网关时每个请求都因 400 报错而失败的问题,该问题为 2.1.275 引入的回归。此版本自 v2.1.275 以来共合并 43 个提交。
一家全球金融科技公司把编码智能体工作负载迁到 Together 的 Dedicated Model Inference,运行 GLM-5.2 并采用多副本 B200、256K 上下文配置。
Microsoft 在 GitHub 新建 microsoft/azure-dev-tools 仓库,面向 Azure 提供开发者体验工具,包括 Canvases 和 skills。仓库目前仅给出这一句说明,未披露具体功能细节、版本号或可用性信息。
Today we're rolling out Projects in Claude Code on desktop and web. A project is one conversation with Claude. It splits the work into threads itself, runs them as parallel cloud sessions, passes context between them, and keeps going when you leave. In beta for select users.
Projects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.
Pragmatic Engineer 发布与 Matt Pocock 的播客访谈,讨论其从声乐教师转型开发者与教育者、创办 Total TypeScript(总销售额超 250 万美元)的经历,以及为 AI 编码 Agent 构建技能的实践。
Steve Yegge 关停了 Gas Town,并承认每月花费数千美元订阅编码智能体,却只做出了 Gas Town 这一个项目。Databricks 向约 3500 名工程师铺开 GPT-6 Astra,其在高复杂度系统设计与长周期任务上"明确"优于 Opus 5 / Sol 5.6,但整体编码支出增加约 60%,公司为此设立专门的 Astra 子预算。
唐杰发文复盘,GLM-5.3-Flash 从首次在国内加速器上运行到承接全部生产流量只用两周,端到端吞吐达 3.2 倍,大量工作由 GLM-5.3 驱动的 Infra Agent 完成。
We’re sharing how GLM-5.3 helped build and optimize the inference infrastructure serving GLM-5.3-Flash. The system went from its first successful run to production readiness in less than two weeks, with end-to-end throughput tripling relative to the initial baseline. The key was dense feedback: local correctness tests, execution traces, microbenchmarks, and end-to-end measurements that enabled targeted hypothesis testing rather than reliance on aggregate performance metrics alone. https://z.ai/blog/glm-built-its-inference-infrastructure
推荐理由:作者复盘了 GLM-5.3 智能体优化推理基础设施的两周过程,提出了可迁移的分层密集反馈方法与工程师角色转变的判断。
GitHub 用 GitHub Copilot app 和 Copilot CLI 把 Copilot agent runtime 从 TypeScript/Node.js 完全重写为超过 80 万行生产级 Rust,AI 智能体编写了大部分代码,跨 128 个 PR 增量合入 main,性能提升数个数量级,主要由一名开发者几个月内完成。
推荐理由:GitHub Copilot 运行时迁移 Rust 的完整复盘,给出智能体并行协作、提示缓存与评审流程的可迁移工程方法。
Today we rolled out Astra to every engineer at Databricks (N=~3500). Some notes that may be helpful to others: 1. Astra unambiguously out performs our previous highest-end models (Opus 5, Sol 5.6) on highly complex tasks, especially those related to high level system design or long range horizontal tasks. 2. Engineers given Astra increased overall coding spend by around 60% compared to baseline. 3. It is not clear Astra meaningfully improves on medium/low complexity coding tasks compared to earlier models. We suspect those tasks are mostly saturated (i.e. perfectly executed) by existing models. 4. We learned above by piloting Astra with around 200 users to gain signal on both quality and cost. We use Unity Gateway to do cohort-based experiments for all new models. 5. We give engineers a sub-budget specific to Astra to encourage them to use Astra selectively on complex tasks while preferring lower cost models for everyday tasks. Our engineers are able to mix-and-match tools and models within their overall budget envelope (we also allow for increased budgets through various mechanisms). These budgets are defined in Unity Gateway and regularly revisited. Note: We do not have robust comparisons of Astra-vs-Fable because we have net yet rolled out Fable widely due to data retention policies.
I’ve got Codex (voice) in CarPlay. And it’s fantastic. I can now build things during road trips, while all 5 kids make a noise and my wife asks why I talk to the AI more than I talk to her. It runs through my Nightblood iOS app, which gives ChatGPT Voice a face, personality and full access to Codex, my Mac and all the native tools.
NVIDIA 技术博客介绍 cuTile Rust(cutile-rs),一个用 Rust 编写 GPU kernel 的 tile 系统,将 Rust 所有权模型扩展到 tile 级 GPU kernel,把可变输出拆分为互不重叠的片段,并在 kernel 启动间保持主机侧所有权契约。
研究者提出共享选择性持久记忆架构,为多轮工具调用的 Agentic LLM 系统保留任务规格、数据 schema、工具配置和输出约束四类可复用上下文,并丢弃会话专属推理轨迹。
Gergely Orosz 采访 OpenAI 七位工程师与工程负责人,报道 Codex 和 ChatGPT Work 如何成为公司内部几乎所有工作的支柱。
掌握 AI 工程技能,你就能主动塑造构建过程:影响要构建什么,并驱动构建循环。以下是实现这一点的关键技能。https://x.com/i/article/2098450134883594240
Cognition 用 GPT-6 Astra 提升 Devin 测试软件并证明其可用的能力,目标是让工程师少审代码、多交付。该能力聚焦于 Devin 对自身工作的验证环节。
Warp now has built-in support for the Grok Build CLI. - Use Warp's rich input for agent prompts, with support for longer pasted prompts and multi-cursor - Use /remote-control to share your agent session to another device - Access the file explorer and code review panels
GitHub Copilot 应用内置 diff、终端和浏览器三个面板,让开发者无需在编辑器、终端和浏览器之间切换即可完成 AI 编码闭环。diff 面板以绿色标注新增、红色标注删除,支持接受改动、留言或让 Copilot 继续修改;终端面板可直接运行项目命令并支持多窗口;浏览器面板可用 Pick & Polish 工具选中元素并让智能体调整。
我们刚刚推出了 CursorBench 4.0! 它包含了新的任务,用于评估模型遵循指令的能力、在具有挑战性的项目上长期工作的表现,并且比之前更难(所以所有模型的得分都更低了)。
Hugging Face 博客介绍 TRL v1.14 的 AsyncGRPOTrainer 新支持:只训练 LoRA 适配器并仅同步适配器(rank-1 仅几 MB)到 vLLM,训练器与 vLLM 副本作为独立 HF Jobs 运行在分开的机器上,通过挂载 Storage Bucket 共享适配器路径,无需 NCCL 或共享本地盘。
Marc Andreessen 发文回顾十五年前“软件吞噬世界”的论断,指出全球前十大公司中科技市值占比已从 31.5% 升至 94.4%,并宣布 a16z 第二次投资 Cognition。
a16z 宣布再次投资 Cognition,押注其 AI 编程智能体 Devin。Devin 在 Cognition 内部编写生产代码的比例已从 13% 升至 90% 以上,并在 Mercedes Benz 将原需八个月的 COBOL 迁移压缩到 8 天,在 Rivian 将测试生成速度提升 10 倍。
Pragmatic Engineer 播客对话 OpenAI Core Products & Platform 负责人、Codex 工程师 Tibo Sottiaux,讨论 Codex 的构建与迭代。
Mistral 帮助一家欧洲能源运营商将 4 万行 Fortran 77 的油藏模拟器迁移到 C++,该代码库没有测试套件和集中文档。团队先搭建数值对齐校验框架,用自定义解析器生成调用树并借助 Vibe CLI 启动上百个智能体补文档,最终采用人工把关的 coder、tester、reviewer 智能体工作流逐模块迁移。
Microsoft 在 GitHub 上线 microsoft/frontier-ghcp-ttt-rvas 新仓库,全称为 Frontier GitHub Copilot Train-The-Trainer Real Value Acceleration Solution(RVAS)。该仓库定位为 GitHub Copilot 的培训师培训与价值加速方案,目前公开信息仅包含仓库名称与全称。
Together AI 发布长文,解析开发者从闭源模型转向开源模型所需的 AI 编码栈,提出由模型、推理、网关与路由、Harness、工具(Skills 与 MCP)组成的 MIGHT 五层框架。
推荐理由:Together AI 把开源编码栈拆为 MIGHT 五层,给出大小模型分工和分层组合的具体实践方法。
Sierra 发布并开源 hyper-𝜏-bench(论文名 𝜏^𝜏-bench),一个衡量模型自主构建客服智能体能力的长程评测。最佳配置 Claude Opus 5(max reasoning)在 Claude Code 中独立通过 23.9% 的保留评测任务,与深度上下文的工程师协作时达 82.2%。