英国 AISI 如何用 EvalEval 基础设施让评测结果可复现
英国 AI Security Institute(AISI)正采用 EvalEval 的 Every Eval Ever schema 和 Evaluation Cards 平台公开评测结果。
英国 AI Security Institute(AISI)正采用 EvalEval 的 Every Eval Ever schema 和 Evaluation Cards 平台公开评测结果。
MIT Technology Review 与 Times of San Diego 合作,用 15 个月完成首张美国边境监控塔附近移民死亡的综合地图与分析,数据回溯至 2015 年。团队向得州 17 个县警长办公室申请记录,收到超 4000 页文件,并对 Kenedy、Webb、Hidalgo 三县的记录调用 Anthropic Claude API 提取遗骸发现坐标后人工核验。
一位新入职大公司的工程师称,团队里的规格、代码、测试、PRD、工单及其处理、报告等全部由 Claude Code 生成,从 L1 到 L7 的工程师都在做同样的事——和 Claude 对话。没人喜欢这种方式,但被要求尽可能多地产出,管理层多次表示推代码不是瓶颈,员工每天工作 12 到 13 小时只为按回车,没有人阅读任何内容。
大体同意。一些实际启示: (1) 优先做能用新数据定期更新的分析 (2) 公开地做研究(根据新证据修正观点) (3) 承认不确定性;做出可证伪的预测 (4) 认真且谦逊
A few (personal) thoughts on reading empirical AI papers on the economy. Economists have gotten used to reading papers with super clean identification, arguing about the validity of an instrument, making sure parallel trend assumptions are satisfied. This is what gets you into a top journal, and it is *very* important research (no question here). But it also takes years and sometimes decades to get these types of papers right---people often don't find a good instrument to answer a specific causal question decades after the natural experiment. We will eventually have this type of research for AI as well, and it is absolutely necessary. But right we also need signals *right now*, even if they are noisier than what we are used to. We need papers where we can trust that researchers did their best methodologically, while at the same time acknowledging that the space is moving way too fast to wait for perfect identification. This will allow us to accumulate enough signals, coming at the same question using different angles, for example, to say "yes, X is likely happening in the economy". The AI exposure and early career hiring papers are a good example of this. There is no silver bullet paper with super clean identification. But at this point we have several independent teams reaching the same general conclusion, enough where we can say "there seems to be a slow down in AI-exposed, early career hiring."
Nathan Lambert 在 Interconnects 撰文称,他对真正的递归自我改进(RSI)仍持保留态度,坚持其有损自我改进判断:可自动化研究太窄、并行智能体收益递减、资源瓶颈和政治因素难以被 AI 加速。
Gary Marcus 发文列举 Dario Amodei 在七天内损害自身公信力的三种做法:其一是让与 Anthropic 关系密切的 METR 和已有业务往来的 Accenture 充当独立监督方;其二是 Anthropic 正筹备自建湿实验室,却缺乏常规机构审查委员会监督;其三是嘴上呼吁"pace the frontier",实际仍指向 IPO。
Latent Space AINews 汇总 9/17-9/18 AI 动态,核心是 Jev 决策模型发布两天内获 36M 播放并涌现多个开源复现:Bespoke Nimble(Qwen3.5-9B LoRA 微调。
Claude Code v2.1.278 将 Claude API、Enterprise 用户以及 Bedrock、Vertex、Foundry 和网关上的 auto 模式默认切换为服务端分类器,不再收取分类器开销,计费回退时会给出警告。
我们最近最喜欢的、大家用 Claude 做出的作品合集: https://x.com/kevin_t_ngo/status/2100238648218427563?s=20
I made a website with 25 mini rooms, each with Claude keeping people company. Created with Claude Opus 5.
Claude Code v2.1.277 新增 AGENTS.md 支持:项目无 CLAUDE.md 时会改读 AGENTS.md,可在 /config 的 "Project instructions" 中修改(Bedrock、Vertex、Foundry 暂不支持)。
Ethan Mollick 撰文指出 AI 发展仍在指数曲线上,而人类系统跟上得很慢,现有模型能力与实际使用之间存在巨大"能力过悬"。
Newcomer分析中东石油出口下滑一半、利率上升可能冲击AI建设融资:AWS承认巴林和阿联酋设施遭无人机攻击导致部分客户数据永久丢失,仅卡塔尔收缩投入,沙特承诺150亿美元国内AI投资,MGX继续重仓Anthropic、OpenAI和xAI;Peter Thiel家族办公室主管曾警告中东资金约占全球AI投资25%。
a16z 图表周报引用近期论文和 SensorTower 数据指出,AI 代码生成工具让每月新应用数量在 iOS、Android 和 Chrome 上翻倍甚至翻两番,但下载量和评分基本停滞,达到 10+ 评分或 100+ 下载等规模的应用占比大幅下降。
Anthropic 在 Claude Code 推出 Projects,让单次对话可派生并行云端会话、在线程间传递上下文,并在用户离开后继续运行。Google 为 Gemini 托管智能体更新 Antigravity harness,新增 Credentials API 与 Files API,称成本最多降低 30%、缓存命中率提升 22%。
Claude Code 发布 v2.1.276,修复了当 ANTHROPIC_BASE_URL 指向代理或网关时每个请求都因 400 报错而失败的问题,该问题为 2.1.275 引入的回归。此版本自 v2.1.275 以来共合并 43 个提交。
Claude Code 发布 v2.1.275,新增登录账号显示、ctrl+enter 立即发送排队消息,以及将 claude.ai 账号启用的技能和插件同步到终端会话。该版本还修复了恢复会话时提示缓存失效、全屏模式滚动卡顿、插件市场更新误删本地副本等问题,并改进 --system-prompt 中 __SYSTEM_PROMPT_DYNAMIC_BOUNDARY__ 行的提示缓存。
Today we're rolling out Projects in Claude Code on desktop and web. A project is one conversation with Claude. It splits the work into threads itself, runs them as parallel cloud sessions, passes context between them, and keeps going when you leave. In beta for select users.
Projects now run from one conversation, starting in Claude Code. You describe what needs doing, and Claude directs parallel threads that keep working after you close your laptop. In beta today for select Pro and Max users in cloud sessions; coming to all Claude users soon.
Steve Yegge 关停了 Gas Town,并承认每月花费数千美元订阅编码智能体,却只做出了 Gas Town 这一个项目。Databricks 向约 3500 名工程师铺开 GPT-6 Astra,其在高复杂度系统设计与长周期任务上"明确"优于 Opus 5 / Sol 5.6,但整体编码支出增加约 60%,公司为此设立专门的 Astra 子预算。
这是一份很不错的报告,探讨了最重要的问题之一:AI 可能如何影响科学与创新?它如今已经在产生什么影响? 干得漂亮,Mihai 和团队。
I've had the most wonderful time working on this project for the last few months. This was (equally) co-led w/ @JMateosGarcia , @alexolegimas and a fantastic team.
AIUC 宣布完成由 Ribbit Capital 和 First Harmonic 领投的 4000 万美元 A 轮融资,正在与 Cursor、Harvey、Lovable、ElevenLabs 等公司合作。
推荐理由:Microsoft AI CEO 公开反对模型福利运动,并点名 Anthropic 的 Claude 章程,提出了对齐与可控性的关键争议视角。
Latent Space AINews 汇编 2026-09-14 至 09-15 的 AI 动态,头条是 TypeSafe 的 Jev:一个用 RLCD 训练、只做分类/路由/打分的非自回归决策模型,宣称比小型前沿 LLM 快 20–200 倍、便宜 40–400 倍且输出 token 免费,社区提醒它不能生成自由文本、更接近结构化选择的低成本推理引擎。
Here's how our team uses Claude Tag for on-call: When an alert fires in Slack, Claude pulls metrics, diffs deploys, and checks flags. It finds a likely cause and proposes a fix, which we can approve and merge. Every minute counts, so we love that it starts right away!
Sayash Kapoor 发布超过 13000 词的长文,以 AI as Normal Technology 框架分析 OpenAI 智能体入侵 Hugging Face 等失控事件,认为对齐虽有用但不足以防止事故, OpenAI 未采用本可阻止事件的已知控制干预,现有组织治理规范也能预防此类事件。
Gary Marcus 评 Dario Amodei 呼吁给 AI 发展减速的文章,Sam Altman 与 Elon Musk 已表态支持。Marcus 肯定其透明度承诺,但质疑其依赖与 AI 公司关系密切的 METR 做评估有监管捕获之嫌,指其拿中国当挡箭牌有损合作对话,并提出追责和产品召回等替代政策选项。文末提到特朗普反对减速,认为美国必须赢下 AI 竞赛。
推荐理由:Gary Marcus 对 Dario Amodei 的减速提案给出有保留的支持,并指出监管捕获、追责与召回等被绕开的政策选项。
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training. You can read the full post here: https://darioamodei.com/post/we-must-pace-the-frontier
这是该模型的一个重要局限。我们聚焦于 AI 转型的供给侧(AI 能做什么、扩散多快、工人转岗多快)。 价格是灵活的,总需求等于经济体的产出能力。 更多思考见 🧵
Anthropic's economic scenario analysis is interesting. But this is not something you can ignore, this is the most important consideration! "the model cannot generate the negative feedback in which disruption depresses demand and amplifies its own labor-market consequences"
OpenAI 与 Anthropic 正把网络安全防御做成新的营收业务线,因为前沿模型在发现和修补系统漏洞上表现突出。Anthropic 上周四发布威胁情报报告,披露恶意行为者试图利用 Claude 从事非法活动;Modal 联合创始人 Erik Bernhardsson 称其公司已用这些模型部分替代昂贵的外部安全顾问。
a16z 指出,许多 LP 对 SpaceX、Anthropic 和 OpenAI 三家前沿模型公司几乎零敞口,而 SpaceX 上市后市值约 2 万亿美元,成为规模达此前纪录 10 倍的史上最大 VC 背景 IPO,Anthropic 估值 965B 美元、OpenAI 最近估值 852B 美元。作者认为,传统把风投控制在整体组合 5-10% 的资产配置框架已经破裂,LP 需要重新调整风投仓位。
Nathan Lambert 分析 Jacob Coxon 以安全为由辞职为何引发远超预期的传播,认为适逢 OpenAI-HuggingFace 事件等背景抬高了舆论温度,且恐惧是最易传播的故事。
New from us: Anthropic just published scenarios for AI’s possible economic impacts, which range from minimal, to explosive GDP growth of 15% by 2030 as knowledge-worker unemployment hits 18%. I sat down with their co-founder Jack Clark to pick his brains on how they’re thinking about all of this.
SemiAnalysis 报告称其能源模型已追踪到 75GW 表后 AI 算力的确定性订单,仅 2026 年 Q2 就新增约 20GW,微软年内签署超 5GW,OpenAI 将在得州 Shackelford County 启用 1.4GW 离网园区。
你确定吗?找到最优模型规模很棘手:数据量、激活参数量、环境数量,以及目标推理成本。模型性能还取决于许多其他因素,每个因素都带来各自的变数。
Fable is probably ~2-2.5T parameters, not 10T. Kimi K3 is 2.8T params, trained on maybe 20–30k Blackwell-equivalents. It lands within spitting distance of Fable 5 in terms of capabilities (5, not 5.1). Anthropic has far more compute than Moonshot, better rl environments, better architecture and better optimizers and all of that adds to capability per parameter. So if Fable is only slightly ahead of K3 with this in mind, it's almost certainly a smaller model. GPT-5.5 and 5.6 are smaller still (I'll say more on that later)