跳到正文

#编码

今日 43 条
今天10月1日周四
  1. HuggingFace Daily Papers(社区热门论文)42

    A2Z GameSpec-Bench:编码智能体能多忠实地按游戏设计文档生成游戏?

    研究者推出 A2Z GameSpec-Bench,用 100 份长篇游戏设计文档(GDD)评测编码智能体的端到端游戏开发忠实度,每份 GDD 被转为含规则、约束与前置依赖的契约,并结合源码检查与智能体生成的测试策略做场景回放和自适应试玩。评测显示当前智能体难以同时满足代码实现与实际游玩中的相互依赖需求;针对具体需求的反馈相比两轮自我修订可将 GDD Fidelity 提升 10.9%。

  2. Google DeepMind:Blog(RSS)77

    Google DeepMind 发布 Gemini 4 Argon 前沿模型

    Google DeepMind 宣布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 向可信网络防御者开放,再逐步扩展至开发者、企业和消费者。

    推荐理由:原文给出定价、1M 输出上限和多项基准成绩,读者可据此评估该模型在编码与防御性网络安全上的实际表现。

  3. Google Blog:AI(RSS)76

    Google 发布 Gemini 4 Argon 前沿模型

    Google 发布新前沿模型 Gemini 4 Argon,先通过 Fairwind Program 面向可信网络防御者开放,价格为每百万输入 token $2、输出 token $10,缓存输入 token 为输入价的 5%。

    推荐理由:官方公告给出定价、输出 token 上限和多个基准分数,读者可以据此评估它在编码与安全防御场景的落点。

  4. Rohan Paul63

    Google 发布 Gemini 4 Argon,Sundar Pichai 称其在复杂工作流、网络防御和软件工程上表现前沿。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  5. 🚨 AI News | TestingCatalog47

    突发 🔥:Google 宣布 Gemini 4 Argon,一款新的前沿模型,面向"跨真实世界软件工程、法律和金融等企业知识工作、以及网络防御的复杂工作流"。 在 DeepSWE v1.1 上取得 77.9% 的分数,创下新 SOTA。在众多基准测试上表现优于 GPT-6 Astra、Opus 5.5 和 Fable 5.1。 即将推出,首先面向付费 API 客户和 Google AI Ultra 订阅用户。 很快!👀

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  6. dex39

    这正是我们讨论杠杆的原因,也是 /show-me 的用途所在 https://www.youtube.com/watch?si=kPO3wH00gA3XQKhV&t=2959&v=-w6RDNBAI0E&feature=youtu.be

    引用Sureffi@Sureffi

    "a spec that is sufficiently detailed to generate code with a reliable degree of quality is roughly the same length and detail as the code itself" ^^ 100% don't believe in that. You should care about the code at the level of abstraction @dexhorthy is describing here. But no way in hell you can't compress it way smaller than the code itself would be (40:1 based on my measurements for a 30k LOC codebase). An LLM holds the priors for pretty much every single convention there is. Along with the cultures from Linus Torvalds to corporate Java. Two things - Understanding the model's priors for your choice of language/framework(s). - You holding those same priors. First screenshot is benchmark results from a few days ago. Second is what the 40:1 compression looks like. YMMW with typescript or python slop.

  7. lauren47

    Grok Bot 现可把编码任务交给 Cursor 云智能体,用户只需在 Slack 中 @ 机器人即可派活,还能创建 Projects 把多个相关智能体归入同一对话。这些云智能体支持 Cursor 上任意模型,且各自拥有独立计算机,使 Bot 更像管理者而非亲自写代码。Bot 还能通过 GitHub 和 Origin 插件管理 PR,并分享构建过程的视频演示。

    引用Grok Bot@bot

    Grok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.

  8. Claude Code:GitHub Releases(RSS)40

    Claude Code v2.1.286 发布

    Claude Code 发布 v2.1.286,为堆叠的权限请求加上"2 of 5"计数,并支持在全屏列表中点击"N more"行跳转。该版本修复了 claude --resume 和 --continue 在并行工具调用后丢失全部轮次、工具或 hook 返回对象/数字/布尔值时触发 API 400 错误,以及云会话因容器在 transcript 加载中被停止而无法唤醒等问题。

  9. Google AI:DEV 作者专属(RSS)59

    Ornith-1.0-9B vs Qwen3.5 vs Gemma4:CPU 本地实测对比

    作者在纯 CPU、32GB 内存的普通笔记本上,用 Ollama 以相同 Q4_K_M 量化对比 Ornith-1.0-9B、其基座模型 Qwen3.5-9B 和 Gemma4-12B,五个任务显示 Ornith 在 JSON 输出上最紧凑(16 token),但 bug 修复在未见过用例上出错,shell 命令与基座同样在含空格文件名上失败。

  10. Google AI:DEV 作者专属(RSS)39

    Sentinel-IR:AI 智能体运营省下数百万成本的非技术指南

    Sentinel-IR 是一种确定性数据压缩层,把代码、API 载荷和文档压缩成超紧凑的中间表示再喂给 LLM,充当上下文窗口的 ZIP 压缩。在 1,366 行的 12-billing-platform 测试文件上,原始 11,635 tokens 被压到 1,332 tokens,节省 88.6%;但 303 tokens(约 34 行)以下的微文件因压缩开销反而更贵。

  11. Google AI:DEV 作者专属(RSS)75

    一次 Agent 重构事故复盘:二十个正确改动如何掩盖了错误的假设

    作者复盘充电站地图去重任务的事故:一个由 Agent 编写、重构后测试全部通过的去重任务,因测试数据自造而未取自真实数据(9269 对重复记录中运营商名称仅 1 对匹配),导致约三分之一注册表站点在 100 米内存在重复显示,重构 78 分钟后被无审阅合并。

    推荐理由:作者以真实去重事故为底,给出从审代码转向审概念与真实数据的可迁移复核清单。

  12. eric zakariasson53

    作者邀请试用 xAI 的 Grok Bot 市场工程类机器人,链接为 https://x.ai/bot/marketplace/engineering。其引用的 Grok Bot 官方内容称,Grok Bot 现在更适合软件开发,机器人可以把编码任务交给 Cursor、通过 GitHub 和 Origin 插件管理 PR,并分享构建内容的视频演示。截图显示工程分类下有多个可选机器人,包括 SWE by Cursor 等。

    引用Grok Bot@bot

    Grok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.

  13. Google AI:DEV 作者专属(RSS)73

    工程师复盘十个月围绕 Claude Code 搭建守护框架的经验与教训

    一位工程师复盘十个月为 Claude Code 逐步搭建 hooks、门禁和守护框架的经历,核心结论是“仪表会撒谎”,提示词只是请求,真正有效的是代码边界。

    推荐理由:作者用十个月的真实失败数据说明提示词为何挡不住模型,并给出可复用的 Claude Code hooks 种子。