跳到正文

#评测/基准

今日 51 条
9月29日周二
  1. Thomas Wolf53

    Deven 将 NanoGPT 训练纪录从 67.6 秒推进到 39.9 秒,通过按 flop 价值跳过计算的稀疏范式实现,包括采样 softmax、稀疏优化器状态与通信、Anvil2 优化器、末段 300 步 EMA 等,并在 8xH100 上将稀疏嵌入参数扩展到 65B,占总收益 25%。Thomas Wolf 转发称其 impressive,并附上 PR 与作者的改动自述 https://github.com/KellerJordan/modded-nanogpt/pull/360 、https://hyperstition.cc/training-nanogpt-in-39-9-seconds 。

    引用Larry Dial@classiclarryd

    New historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds

  2. Diogo Almeida25

    再强调一下,我极度反对基准测试(: 引用推文例外处理——引用推文核心要点:别再搞“Jevbench”了,别再要求公开基准测试,它们完全没抓住Jevons悖论的重点,而且你不敢相信,你珍视的每一个基准测试有多容易被刷分。这是@CompleteSkeptic最惨痛的教训:选对任务胜过一切。

    引用Latent.Space@latentspacepod

    STOP making "Jevbench"es, stop asking for public benchmarks, they completely miss the point of Jev and you won't believe how easy it is to game every benchmark you hold dear This is @CompleteSkeptic's bitterest lesson of all: picking the right task beats everything

  3. Arena.ai49

    OpenAI 的 GPT-6 Luna (Max) 进入 Agent Arena 帕累托前沿,净提升 +1.59%,中位成本仅 $0.05/任务。

    引用Arena.ai@arena

    GPT-6 Luna (Max) by @OpenAI is #23 in Agent Arena with +1.6% net improvement across 8K real-world agentic sessions from our global community of users. Although GPT-6 Luna (Max) did not land on the Agent Arena Pareto frontier, it remains a cost-efficient model. Its $0.05 median cost per task is 94% lower than GPT-6 Sol (Max) at $0.82 and 98% lower than GPT-6 Astra (Max) at $2.59. Its net-improvement score also comes within 0.09 percentage points of #22 GPT 5.5, while costing 91% less than its $0.56 median cost per task. This release is a six-place point-rank move over GPT-5.6 Luna (xHigh), at -0.9% and #29! By signal, GPT-6 Luna’s clearest gains over GPT-5.6 Luna are in: - Confirmed Success: #17 (+4.6%) vs. #33 (-4.6%) - Bash Recovery: #21 (+4.2%) vs. #27 (+2.2%) Congrats to the @OpenAI team on this release!

  4. Diogo Almeida50

    Diogo Almeida 转发引用了他人的 Jev 1.13 红队测试结果:在 DTap 平台上直接滥用下 ASR 为 70.1%,间接提示词注入下为 43.5%,注入可导致数据外泄、删文件等危害;用 Jev 自身作为工具调用的 self-gating 层可显著降低 ASR 并保留大部分效用。作者据此评论,不要只把 Jev 接入高层决策,而应围绕简单原语显式编程想要的行为。

    引用Zhaorun Chen@zrrrr_cn

    Jev is fast at helping you. Turns out, it can also be fast at helping an attacker!! 😱🚨 We red-teamed Jev 1.13 on our DTap (DecodingTrust-Agent Platform) and found a serious safety gap: 70.1% ASR under direct misuse 43.5% ASR under indirect prompt injection In our evaluations, we found that under indirect prompt injection, Jev can follow attacker-injected instructions without blinking an eye, e.g., exfiltrating user data, deleting files, or taking other harmful actions. But we found a much safer way to integrate Jev: use it as a self-gating layer for its own tool calls, significantly reducing ASR while preserving most of its utility. 👇 Read more below

9月28日周一
9月27日周日
9月26日周六
9月25日周五
  1. karminski-牙医43

    阶跃 Step-5-Preview 实测显示输出稳定性突出,后训练扎实,写工程代码很少反复;在"硅基交警"Agent 测试中全程未指挥出事故,后端 Agentic Coding 与"硅基骑手"多轮测试得分 Δ 很小。其 Agent Loop 迭代能力可卡着极限数值优化,在送餐超时规则下主动把超时算成送餐时间+5分钟以赚积分。短板是前端、3D 场景与美学,较上一代有提升但距本代 SOTA 仍有差距。

9月24日周四
9月23日周三
  1. Boris Cherny70

    Boris Cherny 称 Claude Opus 5.5 已成为他近几周的日常主力模型。他让 Opus 5.5 与 Fable 5.1 各自把 HAProxy 从 C 移植到 Rust,两者都通过了几乎全部测试,但 Opus 5.5 用时 9.5 小时,快于 Fable 5.1 的 12 小时,成本低 51%。据其引用的官方介绍,Opus 5.5 是 Claude 5.5 家族首个模型,大多数任务表现达到 Claude Fable 5.1 水平,运行成本比 Opus 5 低 40%。

    引用Claude@claudeai

    Introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.

    推荐理由:作者实测对比 Opus 5.5 与 Fable 5.1 移植 HAProxy 的用时和成本,给出第一手数据供选型参考。

9月22日周二
  1. elsewhere:文章(RSS)68

    阶跃 Step 5 Preview 实测评测:数据可视化与金融分析亮眼,泛化和审美仍有短板

    阶跃发布 Step 5 Preview,总参数量 600B、激活参数 27B,有视觉输入,官方称在 Artificial Analysis 上涨 44 分、单任务成本仅为 Claude Opus 5 的 1/8。作者与友人实测发现其在数据可视化、金融分析上表现不错,但泛化性、领域知识和审美偏弱,思考过程过长导致长任务耗时且易中断,且长上下文下安全指令遵循可被绕过。

  2. eric zakariasson30

    grok 4.7 的一些实际用例! 飞行模拟

    引用aditya@adxtyahq

    GROK 4.7 IS ACTUALLY COMPETING WITH GPT-6 ASTRA. I gave GPT-6 Astra, Grok 4.7, Kimi K3 and Fable 5.1 the same prompt to build a flight simulator Astra was still #1 overall, but Grok 4.7 was surprisingly close Kimi K3 and Fable 5.1 were basically a draw and both produced a much smoother result, while Grok 4.7 was right up there with Astra in terms of overall quality. overall: Astra > Grok ≈ Kimi ≈ Fable GPT-6 Astra finally has some serious competition.

  3. Xiaomi MiMo40

    💗

    引用Design Arena@DesignArena

    BREAKING: MiMo-V2.6-Pro by @XiaomiMiMo lands at #8 overall (#3 open-weight) on Design Arena with an Elo of 1338. This is an impressive 54-point and 22-position increase from MiMo-V2.5-Pro. MiMo-V2.6-Pro also reaches #4 overall in Website (#2 open-weight) and #6 overall in Agentic Frontend Development (#2 open-weight), showing strong performance across both direct generation and agentic coding. This places @XiaomiMiMo's new model among leading models such as GPT-5.6 Sol by @OpenAI, Claude Opus 5 by @Anthropic, and Kimi K3 by @MoonshotAI. Congratulations to the @XiaomiMiMo team on returning to a top-10 placement on Design Arena!

9月19日周六
9月18日周五
  1. Karina56

    Epoch AI 推出 Benchmark Reviews 计划,对 AI 基准进行审计,首批覆盖 15 个基准,其中 4 个 Verified、9 个 Flawed、2 个信息不足。Karina Nguyen 转发并称赞这一举措能激励行业构建真正高质量的基准,并感谢其审计了 PostTrainBench、挖掘出旧的 SimpleQA。

    引用Epoch AI@EpochAIResearch

    Introducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.

9月17日周四
9月16日周三
9月15日周二
9月11日周五
9月8日周二
9月4日周五
9月2日周三