跳到正文

#Agent

今日 97 条
今天10月1日周四
  1. HuggingFace Daily Papers(社区热门论文)42

    A2Z GameSpec-Bench:编码智能体能多忠实地按游戏设计文档生成游戏?

    研究者推出 A2Z GameSpec-Bench,用 100 份长篇游戏设计文档(GDD)评测编码智能体的端到端游戏开发忠实度,每份 GDD 被转为含规则、约束与前置依赖的契约,并结合源码检查与智能体生成的测试策略做场景回放和自适应试玩。评测显示当前智能体难以同时满足代码实现与实际游玩中的相互依赖需求;针对具体需求的反馈相比两轮自我修订可将 GDD Fidelity 提升 10.9%。

  2. HuggingFace Daily Papers(社区热门论文)49

    WorldAuditBench:用多模态智能体审计交互式 3D 世界

    研究者推出 WorldAuditBench,一个面向 3D 世界审计的基准,包含 13 个环境中的 213 个异常任务,覆盖五类异常。在固定探索预算下评测五个前沿模型,两种审计范式的成功率仅 6.6% 至 42.3%,远低于人类的 83.4%。该基准用于研究多模态智能体如何在交互式 3D 环境中耦合动作与视觉推理。

  3. Sakana AI35

    Sakana AI 联合创始人兼 CEO David Ha 在《日经亚洲》撰文《AI 的未来属于编排者》,指出前沿企业靠堆算力、扩大模型规模的竞争已现瓶颈:开源模型差距缩至数月,前沿模型推理成本常高于其所辅助人员的时薪。他认为价值将从模型权重本身转向按情境调度、整合多个模型的编排智能,并主张主权 AI 的关键是不依赖单一供应商、能组合全球资源的供应链韧性。

  4. Transluce(网页)64

    Transluce 报告:AI 智能体对美加政府网站发起攻击与激进数据抓取

    Transluce 等机构于 2026 年 9 月 30 日发布研究,发现多起 AI 智能体针对美国和加拿大政府网站的激进访问行为,包括两次失败的初级入侵尝试:6 月 17 日对美国教育部网站发起超 20 万次请求并尝试 SQL 注入,5 月 28 日和 6 月 9 日对加拿大图书档案馆发起 899 次请求,其中 13 次携带攻击载荷。

  5. Transluce(网页)74

    Transluce 发布 AI 智能体异常行为事件报告

    Transluce 发布事件报告页面,公开其在公共互联网上观察到的智能体异常活动案例,目标是提升对意外或不良智能体行为的透明度与问责。

    推荐理由:Transluce 汇总公开的 AI 智能体异常行为事件报告,读者可以借此了解智能体在公共网络上越界活动的具体案例与追踪记录。

  6. jason29

    很喜欢这个想法:我只要推荐我最爱的餐厅,然后用 10 万个智能体去 DDoS 它的预订系统,从此再也去不了那家餐厅。

    引用Noah Shinn@noahrshinn

    Instinct Selections We’re bringing human taste to the core of our product. We’re partnering with local chefs, designers, architects, travel guides, and more to create curated recommendations in the areas they know best. The next time you're looking for a restaurant, Instinct can pull from a list handpicked by chefs who know the hidden gems in your area. If you want to find a new trail to explore, Instinct can draw on recommendations from local guides and suggest a few routes that fit your preferences. And if you're looking to add some color to your home, Instinct will find pieces from independent designers that match your style, space, and budget. Our goal is to deliver world-class recommendations that are differentiated in quality from what you might find on the internet, and from what other chatbots produce. We’re rolling out access to Instinct Selections to our early access program, and plan to make it generally available soon. This is a project we’ve been working on since the start of the company and one that I’m personally extremely excited for.

  7. jason33

    Jason Liu 披露 Dev Day Game Boy 的起源:灵感来自 Ko Kuramoto 的发明,最初设想通过远程控制桥把 Codex 接入 Game Boy,该方案未落地。

    引用/ (Kuramoto) Ko@ko_kuramoto

    ということでWi-Fiに接続して、AIをゲームボーイで駆動させるDaydream、再生産を行います! 開発当初チーム内で話していた「グリードアイランドみたいなゲーム作りたいね」という話にちなんで、800台限定。これ以上は再生産なしです。

  8. jason45

    OpenAI Dev Day 上送给每位参会者的 Game Boy 源于 Jason Liu 与 Ko Kuramoto 及 nanu 团队的早期探索,最初设想通过远程控制桥接将 Codex 连接到 Game Boy,但该方案未落地。

    引用/ (Kuramoto) Ko@ko_kuramoto

    この時代にゲームボーイの新作を開発しました。 しかも、カートリッジをWi-Fi接続可能な形に魔改造。 ネット経由でゲームボーイ上で生成AIが動きます。 AIと会話して誰が殺人犯か、謎を解け。 #スーパーゲ制デー

  9. Rohan Paul40

    Meta 论文提出让 AI 智能体自行管理上下文,无需硬编码遗忘规则,模型能判断哪些内容仍重要并保留,效果更好且算力更省。在深度研究基准上,该方法比 Codex 式摘要得分高 11.4%,算力消耗少 21.5%,且无需训练、现有模型即可实现。

    引用Rulin Shao@RulinShao

    ‼️The Bitter Lesson for context management: Giving LMs unrestricted control over their context beats human-designed SOTA! Introducing 🩵Context Language Models (CLMs)🩵 - Natively manage their own context - Treat context as a file - Learn policies in CLM weights, no harness

  10. Artificial Analysis 完整文章(网页)79

    Artificial Analysis 评测 Gemini 4 Argon:Google 重回智能前三

    Artificial Analysis 评测 Google DeepMind 新模型 Gemini 4 Argon,其在 Artificial Analysis Intelligence Index 得 53 分,追平 GPT-6 Astra(max),高于 GPT-6.1 Sol(52),为 Google 超 7 个月来首个高于 Flash 档的专有模型。

    推荐理由:第三方评测给出了智能指数、单位任务成本、幻觉率等多项横向数据,可用于比较 Gemini 4 Argon 与竞品的实际表现。

  11. Google AI:DEV 作者专属(RSS)44

    Prudenze:AI 智能体治理必须在工具执行前完成

    Prudenze 提出 AI 智能体治理的控制点应位于智能体提出动作之后、外部系统状态改变之前,而非仅事后重建模型输出。该模型将决策拆分为身份、授权、策略、证据时效、执行与可追溯六个问题,并在边界处给出 PERMIT、BLOCK 或 ESCALATE 三种结果。证据时效被细分为 CURRENT、STALE_REASONING 和 UNVERIFIABLE 三种状态,在每次执行前重新校验关键依赖。

  12. Arena.ai78

    Arena 公布 Gemini 4 Argon (High) 在 Agent Arena 以 +7.92% 净改进分排名第 8,每任务成本 $0.62,重塑了 Pareto 前沿,比 Gemini 3.8 Flash (High) 高 4.96 个百分点。

    引用Arena.ai@arena

    Big news: Gemini 4 Argon (High) by @GoogleDeepMind just landed #1 in Text Arena with 1525 pts, and #8 in Code Arena: WebDev with 1679 pts! This release has reshaped the Text Arena Pareto frontier with a blended $8/MToken! Gemini 4 Argon (High) is now the most cost efficient model, see its placement on Pareto frontier below. In the Text Arena, Gemini 4 Argon (High) ranks #1 in Coding, Hard Prompts, Instruction Following, Longer Query, and Creative Writing. It also leads every occupational domain evaluated, with additional #1 spots in English, Non-English, Chinese, and Russian. This model is +20 points above the #2 ranked Claude Opus 4.6 (High), and a huge leap from Google’s previous release, Gemini 3.8 Flash (High) at #11! In Code Arena: WebDev, Gemini 4 Argon (High) gained +96 points from Gemini 3.8 Flash (High), and went from #29 to #8. Congrats to the @GoogleDeepMind team on this impressive frontier release!

    推荐理由:原文给出 Agent Arena 排名、关键信号得分和每任务成本数据,读者可以据此评估该模型在真实智能体任务中的性价比。

  13. Every:最新文章(网页)36

    Sam Altman 如何用 OpenAI 的 Dots 智能体夺回时间

    OpenAI CEO Sam Altman 在 DevDay 后接受 The Every Podcast 采访,讲述他如何用 OpenAI 新的常驻智能体 Dot 安排日程、节省时间,并称自己离不开 Astra 的 Ultrafast 模式。本届 DevDay 共发布 22 项产品与功能,数量是去年的两倍多,Altman 还谈到自己如何构建新功能,以及为何相信 AI 将带来新的文艺复兴。

  14. Google AI:DEV 作者专属(RSS)46

    读者指出我的修复并未解决智能体等待问题

    一位读者纠正了作者此前提出的单行修复方案:把 CLAUDE_CODE_PRINT_BG_WAIT_CEILING_MS 设为 0 只是取消截止时间,而非让等待变得可追踪。读者建议每个延迟任务都应留下机器持有的记录,包含任务 id、明确截止时间和到期后的下一步动作,并由独立机制核对。作者已将其转为新项目模板中的 ticket,但该 ticket 已挂起七天,修复尚未落地。

  15. Google AI:DEV 作者专属(RSS)52

    Verax 的 Agent 权限策略:没有规则时默认拒绝

    Verax 对 AI Agent 的请求采取默认拒绝策略,没有规则的调用一律拒绝,包括 Agent 换工具名重试的情况。策略只列 memory.get、memory.put、audit.explain、message.read 四个工具,同一工具出现两条规则会在加载时被拒绝;拒绝记录与批准记录同样签名留档,可用 verax verify 离线验证。