跳到正文

#评测/基准

今日 0 条
9月30日周三
  1. ARC Prize:官方博客35

    ARC Prize 社区通讯:用 ARC-AGI 赢下黑客松

    ARC Prize 社区通讯披露,Aditya Advani 团队凭 ARC-AGI-Pub 项目在旧金山 AGI House 黑客松拿下第三名。ARC-AGI-Pub 排行榜已上线,官方为其设立 15 万美元验证基金,该榜不计入 ARC Prize 2024 且无奖金。Kaggle 参赛者增至 6,539 人、434 支队伍,每日提交上限从 5 次降至 3 次以抑制过拟合。

  2. ARC Prize:官方博客30

    ARC Prize 2024 美国大学巡回宣讲启动

    ARC Prize 启动 2024 美国大学巡回宣讲,将走访斯坦福、MIT、UC Berkeley 等十余所高校,联合 AI 学生与研究者推进开放 AGI 进展。联合创始人 Mike Knoop 和 François Chollet 将现场讲解 ARC-AGI 的历史、动机、技术路径与未来方向。10 月 24 日另设一场面向公众的线上活动,内容与校内场次相同。

  3. ARC Prize:官方博客34

    ARC Prize 2024 代码提交关闭,进入排行榜验证阶段

    ARC Prize 2024 Kaggle 竞赛代码提交已于 11 月 10 日截止,共 1,451 支队伍提交 19,423 份代码方案,Kaggle 与 ARC Prize 团队已启动排行榜验证。论文提交截止 11 月 12 日,开源截止 11 月 24 日,仅开源方案有资格获奖并进入排行榜,获奖者将于 12 月 6 日公布。ARC Prize 2025 竞赛计划于明年第一季度启动。

  4. ARC Prize:官方博客15

    ARC Prize Foundation:为 AGI 打造人类校准基准

    ARC Prize Foundation 是一个致力于加速 AGI 发展的非营利组织,主张真正的 AGI 不能只靠扩大现有模型规模,而需要转向具备流体智能的系统。该组织通过创建和策划人类校准的基准来衡量人类与 AI 之间的能力差距,并推动研究者探索超越模式匹配与记忆的方法。其团队由 ARC-AGI 创造者 François Chollet 等人组成,并维护开源协作生态。

9月29日周二
  1. Thomas Wolf53

    Deven 将 NanoGPT 训练纪录从 67.6 秒推进到 39.9 秒,通过按 flop 价值跳过计算的稀疏范式实现,包括采样 softmax、稀疏优化器状态与通信、Anvil2 优化器、末段 300 步 EMA 等,并在 8xH100 上将稀疏嵌入参数扩展到 65B,占总收益 25%。Thomas Wolf 转发称其 impressive,并附上 PR 与作者的改动自述 https://github.com/KellerJordan/modded-nanogpt/pull/360 、https://hyperstition.cc/training-nanogpt-in-39-9-seconds 。

    引用Larry Dial@classiclarryd

    New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds

9月26日周六
9月24日周四
9月22日周二
9月18日周五
  1. Karina56

    Epoch AI 推出 Benchmark Reviews 计划,对 AI 基准进行审计,首批覆盖 15 个基准,其中 4 个 Verified、9 个 Flawed、2 个信息不足。Karina Nguyen 转发并称赞这一举措能激励行业构建真正高质量的基准,并感谢其审计了 PostTrainBench、挖掘出旧的 SimpleQA。

    引用Epoch AI@EpochAIResearch

    Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.

9月8日周二
8月6日周四
6月15日周一