跳到正文

#Agent

今日 112 条
今天10月1日周四
  1. Artificial Analysis 完整文章(网页)79

    Artificial Analysis 评测 Gemini 4 Argon:Google 重回智能前三

    Artificial Analysis 评测 Google DeepMind 新模型 Gemini 4 Argon,其在 Artificial Analysis Intelligence Index 得 53 分,追平 GPT-6 Astra(max),高于 GPT-6.1 Sol(52),为 Google 超 7 个月来首个高于 Flash 档的专有模型。

    推荐理由:第三方评测给出了智能指数、单位任务成本、幻觉率等多项横向数据,可用于比较 Gemini 4 Argon 与竞品的实际表现。

  2. Google AI:DEV 作者专属(RSS)44

    Prudenze:AI 智能体治理必须在工具执行前完成

    Prudenze 提出 AI 智能体治理的控制点应位于智能体提出动作之后、外部系统状态改变之前,而非仅事后重建模型输出。该模型将决策拆分为身份、授权、策略、证据时效、执行与可追溯六个问题,并在边界处给出 PERMIT、BLOCK 或 ESCALATE 三种结果。证据时效被细分为 CURRENT、STALE_REASONING 和 UNVERIFIABLE 三种状态,在每次执行前重新校验关键依赖。

  3. Arena.ai78

    Arena 公布 Gemini 4 Argon (High) 在 Agent Arena 以 +7.92% 净改进分排名第 8,每任务成本 $0.62,重塑了 Pareto 前沿,比 Gemini 3.8 Flash (High) 高 4.96 个百分点。

    引用Arena.ai@arena

    Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts! This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below. In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian. This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11! In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8. Congrats to the @GoogleDeepMind team on this impressive frontier release!

    推荐理由:原文给出 Agent Arena 排名、关键信号得分和每任务成本数据,读者可以据此评估该模型在真实智能体任务中的性价比。

  4. Every:最新文章(网页)36

    Sam Altman 如何用 OpenAI 的 Dots 智能体夺回时间

    OpenAI CEO Sam Altman 在 DevDay 后接受 The Every Podcast 采访,讲述他如何用 OpenAI 新的常驻智能体 Dot 安排日程、节省时间,并称自己离不开 Astra 的 Ultrafast 模式。本届 DevDay 共发布 22 项产品与功能,数量是去年的两倍多,Altman 还谈到自己如何构建新功能,以及为何相信 AI 将带来新的文艺复兴。

  5. Google AI:DEV 作者专属(RSS)46

    读者指出我的修复并未解决智能体等待问题

    一位读者纠正了作者此前提出的单行修复方案:把 CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS 设为 0 只是取消截止时间,而非让等待变得可追踪。读者建议每个延迟任务都应留下机器持有的记录,包含任务 id、明确截止时间和到期后的下一步动作,并由独立机制核对。作者已将其转为新项目模板中的 ticket,但该 ticket 已挂起七天,修复尚未落地。

  6. Google AI:DEV 作者专属(RSS)52

    Verax 的 Agent 权限策略:没有规则时默认拒绝

    Verax 对 AI Agent 的请求采取默认拒绝策略,没有规则的调用一律拒绝,包括 Agent 换工具名重试的情况。策略只列 memory.get、memory.put、audit.explain、message.read 四个工具,同一工具出现两条规则会在加载时被拒绝;拒绝记录与批准记录同样签名留档,可用 verax verify 离线验证。

  7. Rohan Paul63

    Google 发布 Gemini 4 Argon,Sundar Pichai 称其在复杂工作流、网络防御和软件工程上表现前沿。

    引用Sundar Pichai@sundarpichai

    Lots of discussion out there about our next model(!), so I wanted to give an early look as soon as possible. Introducing Gemini 4 Argon! It shows frontier performance in complex workflows, cyber defense and software engineering. Teams are using it extensively at Google, from coding to quantum computing, great feedback. Here’s a look at the benchmarks:

  8. Bloomberg:Technology(RSS)32

    Robinhood CEO 称新 AI 智能体应用安全可靠

    Robinhood 董事长兼 CEO Vlad Tenev 在 Houston 峰会上发布公司新的 AI 智能体应用,并称其安全可靠。他表示要让个人交易者获得对冲基金级别的工具,包括跨资产类别 24/7 交易、用户睡眠时仍可运行的自主“agent loops”,以及卫星影像和区块链分析等专用数据源。

  9. TechCrunch:AI(RSS)50

    OpenAI 推出 Decisions API,对标 TypeSafe AI 的 Jev 决策模型

    OpenAI 在 Dev Day 上发布 Decisions API,可为 Luna 模型预设选项并输出概率,功能类似 TypeSafe AI 本月初推出的软件自动化模型 Jev。该 API 目前为限量预览,OpenAI 称其能保持图像理解、多语言支持和安全防护的同时实现极快速度。有开发者用 Jev 监控 AI 智能体行为,每次动作监控成本仅 2.94 美元,而前沿 LLM 需 372 美元。

  10. METR:Blog(网页)70

    METR 主席 Chris Painter 就 OpenAI / Hugging Face 智能体事件向美国参议院作证

    2026年9月30日,METR 主席 Chris Painter 在美国参议院国土安全委员会小组听证会上就 OpenAI / Hugging Face 事件作证。

    推荐理由:这是当事机构负责人在参议院听证会的完整证词,以第一手视角拆解了 OpenAI / Hugging Face 事件的三要素框架,并给出行业共性观察。

  11. lauren47

    Grok Bot 现可把编码任务交给 Cursor 云智能体,用户只需在 Slack 中 @ 机器人即可派活,还能创建 Projects 把多个相关智能体归入同一对话。这些云智能体支持 Cursor 上任意模型,且各自拥有独立计算机,使 Bot 更像管理者而非亲自写代码。Bot 还能通过 GitHub 和 Origin 插件管理 PR,并分享构建过程的视频演示。

    引用Grok Bot@bot

    Grok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.

  12. Google AI:DEV 作者专属(RSS)59

    Ornith-1.0-9B vs Qwen3.5 vs Gemma4:CPU 本地实测对比

    作者在纯 CPU、32GB 内存的普通笔记本上,用 Ollama 以相同 Q4_K_M 量化对比 Ornith-1.0-9B、其基座模型 Qwen3.5-9B 和 Gemma4-12B,五个任务显示 Ornith 在 JSON 输出上最紧凑(16 token),但 bug 修复在未见过用例上出错,shell 命令与基座同样在含空格文件名上失败。

  13. Google AI:DEV 作者专属(RSS)39

    Sentinel-IR:AI 智能体运营省下数百万成本的非技术指南

    Sentinel-IR 是一种确定性数据压缩层,把代码、API 载荷和文档压缩成超紧凑的中间表示再喂给 LLM,充当上下文窗口的 ZIP 压缩。在 1,366 行的 12-billing-platform 测试文件上,原始 11,635 tokens 被压到 1,332 tokens,节省 88.6%;但 303 tokens(约 34 行)以下的微文件因压缩开销反而更贵。

  14. Google AI:DEV 作者专属(RSS)40

    我无法为代码辩护:i.c.stars 第四周的四份文档与产品规则设计

    i.c.stars 第四周产出四份非代码文档:缺陷日志、同行评审、三份 runbook 和一份 run of show,核心都是让工作能经受交接。作者在同行评审中复现对方八项发现中的六项,指出"擅长找 bug、弱于记录 bug"是常态。其产品贡献是主张检索系统只允许两种行为:附文档、章节和生效日期作答,或停止并转交人工,拒绝不是兜底而是同等功能。

  15. Google AI:DEV 作者专属(RSS)75

    一次 Agent 重构事故复盘:二十个正确改动如何掩盖了错误的假设

    作者复盘充电站地图去重任务的事故:一个由 Agent 编写、重构后测试全部通过的去重任务,因测试数据自造而未取自真实数据(9269 对重复记录中运营商名称仅 1 对匹配),导致约三分之一注册表站点在 100 米内存在重复显示,重构 78 分钟后被无审阅合并。

    推荐理由:作者以真实去重事故为底,给出从审代码转向审概念与真实数据的可迁移复核清单。

  16. eric zakariasson53

    作者邀请试用 xAI 的 Grok Bot 市场工程类机器人,链接为 https://x.ai/bot/marketplace/engineering。其引用的 Grok Bot 官方内容称,Grok Bot 现在更适合软件开发,机器人可以把编码任务交给 Cursor、通过 GitHub 和 Origin 插件管理 PR,并分享构建内容的视频演示。截图显示工程分类下有多个可选机器人,包括 SWE by Cursor 等。

    引用Grok Bot@bot

    Grok Bot is now more powerful for building software. Bots can hand off coding tasks to Cursor, manage your PRs with GitHub and Origin plugins, and share video demos of what they build.

  17. Aravind Srinivas69

    Perplexity CEO Aravind Srinivas 宣布向所有人开放 Computer 的邮箱任务功能,无需 Perplexity 账号,限时免费执行所有通过邮箱委派的任务。用户将邮件发送、转发或抄送到 computer@perplexity.com,智能体会在后台完成任务并保留邮件上下文;每个任务以正常会话形式运行,可在网页和移动端查看,审计记录与应用内任务一致。

    引用Perplexity@perplexity_ai

    Computer now works in email. Send, forward, or cc computer@perplexity.com on any thread. Every email task runs as a normal session in Computer, viewable on web and mobile, with the same audit trail as any task in the app.

    推荐理由:Perplexity 开放无需账号的邮箱入口并限时免费,读者可以据此评估把任务转给 Computer 智能体执行的可行性。

  18. MiniMax (official)43

    基于 MiniMax H3 构建,@Creatify_Labs 的 Boreal-H3 是一款专为广告优化的视频模型,在更精准遵循创意简报的同时,保持产品和角色的一致性。 很高兴看到 MiniMax H3 成为更多面向特定行业的前沿模型的基础!✨

    引用Creatify Labs@Creatify_Labs

    Introducing Boreal-H3 — a video model built for ads and our next step toward recursive self-improvement in video generation. A good-looking video isn’t enough. The product has to stay the same. The actor has to stay the same. The label has to be right. And the action in the brief actually has to happen. So we post-trained MiniMax H3 specifically for advertising. But this isn’t a one-off SFT or LoRA fine-tune. We built a closed-loop system that learns what to improve next. Human-calibrated evaluation diagnoses failures and guides the next intervention: targeted data collection, reinforcement learning, or inference optimization. When the feedback is unreliable, we revise the evaluator or reward—not just the generator. Every experiment feeds into shared memory, informing the next training decision. The model improves, and so does the process that produces its successor. The results: → 85.3% reference fidelity — highest among the frontier video generation models we evaluated → Brief success: 28% → 50% → Identity match: 83% → 94% → Visible defects per clip: down 70% → Generation time and estimated cost: down 20% Boreal-H3 doesn’t just make better-looking video. It makes more usable ads. Credit to the @MiniMax_AI team for the foundation we’re building on. This launch is a checkpoint, not the finish line. We’re building more than a better video model. We’re building a system that learns how to make the next one better.

  19. The Decoder:AI News(RSS)68

    Google 在 Gemini 中全球推出 Skills,取代 Gems

    Google 在 Gemini 聊天中全球推出 Skills,一种可由 AI 协助编写和完善的任务详细提示词格式,取代 Gems。用户将常用指令保存为 Skill,输入 "/" 加名称调用,Gemini 还能从历史对话生成 Skills、自动运行匹配的 Skill、链式组合多个 Skills,并支持文本文档、PDF 和图片作参考材料。