Epoch AI 面向机构提供定制 AI 研究与专家咨询服务
Epoch AI 开放定制研究合作,可承接 AI 趋势报告、模型评测、数据收集与基准开发等项目,曾与 OpenAI 合作开发 FrontierMath、与 Google DeepMind 合著 Rosetta Stone 论文、与 EPRI 联合发布 AI 电力需求研究。
Epoch AI 开放定制研究合作,可承接 AI 趋势报告、模型评测、数据收集与基准开发等项目,曾与 OpenAI 合作开发 FrontierMath、与 Google DeepMind 合著 Rosetta Stone 论文、与 EPRI 联合发布 AI 电力需求研究。
Epoch AI 数据显示,自 2026 年 1 月以来,最强开源权重模型平均落后前沿闭源模型约 4 个月,即 8 个 ECI 分。该机构此前测得这一差距在 2025 年 10 月约为 3 个月。其 IKEA 家具组装基准上,AI 模型得分 10 个月内从 28% 升至 80%,开源权重模型落后闭源约 7 个月。
ARC Prize 社区通讯披露,Aditya Advani 团队凭 ARC-AGI-Pub 项目在旧金山 AGI House 黑客松拿下第三名。ARC-AGI-Pub 排行榜已上线,官方为其设立 15 万美元验证基金,该榜不计入 ARC Prize 2024 且无奖金。Kaggle 参赛者增至 6,539 人、434 支队伍,每日提交上限从 5 次降至 3 次以抑制过拟合。
ARC Prize 社区通讯宣布 Team MindsAI(Jack Cole、Michael Hodel、Mohammed Osman)以 41% 的新 SOTA 成绩首次突破 ARC-AGI 40% 大关。
ARC Prize 竞赛今年还剩 86 天,MindsAI 仍以 43% 位居榜首。François Chollet 在 AGI-24 大会上就 AGI 的定义、测量与构建发表演讲,Vladimir Iglovikov 团队则在 Llamathon 活动上用 Llama 模型尝试 ARC Prize 解法。
ARC Prize 推出 Daily Puzzle,每天 UTC 12 点从 ARC-AGI 任务中选出一道作为当日谜题,玩家可分享解题用时与尝试次数。MindsAI 团队在突破 40% 门槛三天后把分数提升到 43%,目前已有超 700 支队伍、近 4000 次提交。
ARC Prize 2024 赛程过半,ARC-AGI 世界纪录提升 13%,达到 46%,由 Jack Cole 与团队 MindsAI 连续第三次刷新。Kaggle 参赛者已超 11,000 人,目前有 4 支队伍得分超过 30%。
ARC Prize 发布 2024 竞赛三个月更新:ARC-AGI 基准仍未被攻破,SOTA 从 34% 升至 46%(MindsAI 通过测试时微调语言模型达成),Kaggle 上有 921 支队伍、7,368 次提交。
ARC Prize 启动 2024 美国大学巡回宣讲,将走访斯坦福、MIT、UC Berkeley 等十余所高校,联合 AI 学生与研究者推进开放 AGI 进展。联合创始人 Mike Knoop 和 François Chollet 将现场讲解 ARC-AGI 的历史、动机、技术路径与未来方向。10 月 24 日另设一场面向公众的线上活动,内容与校内场次相同。
ARC Prize 团队本月走访了 13 所美国大学,与超过 1,500 名学生和教授交流 AGI 进展并推广 ARC Prize。团队还将于明天举办线上活动,由 ARC-AGI 创作者 François Chollet 讲解可能攻克 ARC 并赢得大奖的技术路径。
ARC Prize 2024 Kaggle 竞赛代码提交已于 11 月 10 日截止,共 1,451 支队伍提交 19,423 份代码方案,Kaggle 与 ARC Prize 团队已启动排行榜验证。论文提交截止 11 月 12 日,开源截止 11 月 24 日,仅开源方案有资格获奖并进入排行榜,获奖者将于 12 月 6 日公布。ARC Prize 2025 竞赛计划于明年第一季度启动。
ARC Prize Foundation 宣布推出 ARC Prize Verified 计划,通过官方隐藏测试集认证模型在 ARC-AGI 上的成绩,通过者可进入官方榜单并获得认证徽章。
ARC Prize 2026 提供 200 万美元奖金、设 3 条赛道,面向开源 AGI 进展。同期 ARC Prize 2025 为 100 万美元竞赛,目标是开源 ARC-AGI-2 的解法;ARC-AGI-3 已开放开发者预览,供开发者构建并测试智能体。
Epoch AI 估算华为 2026 年 AI 算力产量不足 Nvidia 的 4%,若无法获得外国存储芯片,2028 年份额可能降至约 1%。OpenAI 的 GPT-6 Astra 以 166 分登顶 Epoch Capabilities Index,Math-ECI 达 170 创纪录,但 SWE-ECI 164 仍落后 Claude Fable 5.1 的 167。
Epoch AI 数据显示,OpenAI 的 GPT-6 Astra 以 166 分登顶 Epoch Capabilities Index,Math-ECI 达 170 创纪录,但 SWE-ECI 164 仍落后 Claude Fable 5.1 的 167。
Epoch AI 发布多项研究:GPT-6 Astra 以 ECI 166 分登顶能力指数,Math-ECI 170 创纪录,但 SWE-ECI 164 仍落后 Claude Fable 5.1 的 167。
Artificial Analysis 推出 Cyber Index,联合 Collinear AI、Vercel、Berkeley RDI 三家合作伙伴整合 CWE-Bench-AA、DeepsecBench-AA、CyberGym-E2E-AA 三项网络安全评测,覆盖从扫描代码漏洞到复现崩溃并打补丁的完整防御流程。
ARC Prize 公布 ARC-AGI 社区排行榜,Tycho 在 ARC-AGI-3 Public Demo 上取得 100.0% 成绩、成本 $2,986,位列榜首。
ARC Prize 发布发展时间线,回顾 ARC-AGI 从 2019 年 François Chollet 论文《On the Measure of Intelligence》到 2026 年推出 ARC-AGI-3 的历程。
ARC Prize Foundation 是一个致力于加速 AGI 发展的非营利组织,主张真正的 AGI 不能只靠扩大现有模型规模,而需要转向具备流体智能的系统。该组织通过创建和策划人类校准的基准来衡量人类与 AI 之间的能力差距,并推动研究者探索超越模式匹配与记忆的方法。其团队由 ARC-AGI 创造者 François Chollet 等人组成,并维护开源协作生态。
ARC Prize Foundation 发起捐赠,用于扩展其工作并打造更具影响力的 AI 基准,捐款可抵税。该机构称 ARC-AGI 是唯一针对"人类能解、AI 尚不能解"这一关键缺口的基准。Google AI Studio 产品负责人 Logan Kilpatrick 表示已个人支持该基金会,以更好判断通往 AGI 的进展。
New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds
SemiAnalysis 拆解 Intel Panther Lake,解析 Intel 18A 首个商用背面供电(PowerVia)、RibbonFET 四层纳米片和 Foveros-S 封装。实测显示 18A 计算逻辑密度与 TSMC N3E GPU 逻辑相近,但不及 N3P、N2 和 Samsung SF2 的峰值密度;NPU 5 面积缩小 36.9%,GPU 高端版本仍用 TSMC N3E。
SemiAnalysis 发布 ClusterMAX 3.0,覆盖 77 家 neocloud 供应商、323 家市场全景,并访谈超 200 名终端用户。Nebius 加入 CoreWeave 进入铂金级,Google Cloud 升入金级,Azure 降至银级,全球仅 19 家获得奖章评级,另新增介于铜级与不合格之间的 Participation Ribbon 级。
英国 AI Security Institute(AISI)正采用 EvalEval 的 Every Eval Ever schema 和 Evaluation Cards 平台公开评测结果。
Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
SemiAnalysis 在 InferenceX 官方预览中发布首个 TPUv7 Ironwood 第三方推理结果,FP8 聚合服务下每美元性能最高比 B200/B300 好 50%,20 tok/s/user 时每百万 token 成本 0.181 美元,低于 B200 的 0.222 美元和 B300 的 0.276 美元。
Chips and Cheese 评析 NVIDIA 45 页 Vera 白皮书,认为 Olympus 核心(10 宽解码、值预测、88 核单die、1.2 TB/s LPDDR5X)确实强劲。
英国 AI Security Institute 对齐团队与 Timaeus 联合成立非营利研究机构 Sequent,认为对齐研究"尚未走上正轨",计划两年内扩至 40-80 名全职员工,初期募资 1 亿至 1.5 亿美元,研究方向涵盖可扩展监督、学习理论、启发式论证、博弈论与人格。