Claude suddenly stopped cheating.
#评测/基准
#评测/基准
今日 51 条
Thomas Wolf@Thom_WolfAI 评分4444引用Lukas Petersson@lukaspet
Thomas Wolf@Thom_WolfAI 评分5353引用Larry Dial@classiclarrydNew historic NanoGPT record at 39.9s (-27.7s) from @DevenPzak , obliterating the prior record of 67.6s! This record introduces a new paradigm of thinking to NanoGPT: instead of optimizing matmuls or adding more expressive operations, optimize at the individual flop level with incredibly clever engineering and ML judgement. If a flop is low value on a particular step, skip it. Specifically: -(~8s) Sampled softmax. If a token doesn’t appear in a batch, skip its lm_head fwd/bwd some fraction of the time. -Sparse values. Only run an optimizer step for ngram embeddings that occurred in the batch. Set beta1 to zero to enable this. Beta2 is applied retroactively when the row is later used. -Sparse updates. Only update ngram and value embeddings once every 4 steps instead of once every 2. -Sparse communication. Shard the n-gram table across GPUs, and only pass the rows receiving updates on each step. -Sparse optimizer states. For the n-gram table, reduce from 2 floats in Adam optimizer per param, to 1 float per 768 params. -Hand-rolled flash attention for 64 dim heads. There are several additions that add accuracy too: -(~4s) EMA during last 300 steps, combined with lifting final_lr to 0.3 instead of 0.15. -(~1s) A new optimizer, Anvil2, which expands muon via a second tracked momentum buffer, improves the ortho coefficients, and modifies the cautious weight decay application. -A couple additional dynamic skip connections in the network. The most striking consequence of the ‘flop aware paradigm’ is you can grow parameters arbitrarily large, only limited by the available memory, since you can selectively choose how to expend flops on those parameters on each step. NanoGPT has kept active parameters below 124M, but total is unbounded, and has grown to 640M through embedding sparsity over the last year. This PR takes that to its logical conclusion on the 8xH100, scaling up to 65B sparse embedding parameters, which accounts for 25% of the PR’s gains. At frontier scale, where one is not bounded by an 8xH100, one could imagine where this paradigm could lead. https://github.com/KellerJordan/modded-nanogpt/pull/360 As this was a very notable PR, I spoke with Deven for an hour to learn how he did it. Here’s his story on the changes: https://hyperstition.cc/training-nanogpt-in-39-9-seconds
Diogo Almeida@CompleteSkepticAI 评分2525引用Latent.Space@latentspacepodSTOP making "Jevbench"es, stop asking for public benchmarks, they completely miss the point of Jev and you won't believe how easy it is to game every benchmark you hold dear This is @CompleteSkeptic's bitterest lesson of all: picking the right task beats everything
Arena.ai@arenaAI 评分4949OpenAI 的 GPT-6 Luna (Max) 进入 Agent Arena 帕累托前沿,净提升 +1.59%,中位成本仅 $0.05/任务。
引用Arena.ai@arenaGPT-6 Luna (Max) by @OpenAI is #23 in Agent Arena with +1.6% net improvement across 8K real-world agentic sessions from our global community of users. Although GPT-6 Luna (Max) did not land on the Agent Arena Pareto frontier, it remains a cost-efficient model. Its $0.05 median cost per task is 94% lower than GPT-6 Sol (Max) at $0.82 and 98% lower than GPT-6 Astra (Max) at $2.59. Its net-improvement score also comes within 0.09 percentage points of #22 GPT 5.5, while costing 91% less than its $0.56 median cost per task. This release is a six-place point-rank move over GPT-5.6 Luna (xHigh), at -0.9% and #29! By signal, GPT-6 Luna’s clearest gains over GPT-5.6 Luna are in: - Confirmed Success: #17 (+4.6%) vs. #33 (-4.6%) - Bash Recovery: #21 (+4.2%) vs. #27 (+2.2%) Congrats to the @OpenAI team on this release!
Arena.ai@arenaAI 评分5050
Simon Willison 博客AI 评分7474 Anthropic 发布 Claude Sonnet 5.5,速度提升 30% 以上
Anthropic 发布新模型 Claude Sonnet 5.5,官方称其运行速度快 30% 以上,多数工作成本最多降低 30%,定价与 Sonnet 5 相同但在各项基准上均优于 Sonnet 5。
Diogo Almeida@CompleteSkepticAI 评分5050引用Zhaorun Chen@zrrrr_cnJev is fast at helping you. Turns out, it can also be fast at helping an attacker!! 😱🚨 We red-teamed Jev 1.13 on our DTap (DecodingTrust-Agent Platform) and found a serious safety gap: 70.1% ASR under direct misuse 43.5% ASR under indirect prompt injection In our evaluations, we found that under indirect prompt injection, Jev can follow attacker-injected instructions without blinking an eye, e.g., exfiltrating user data, deleting files, or taking other harmful actions. But we found a much safer way to integrate Jev: use it as a self-gating layer for its own tool calls, significantly reducing ASR while preserving most of its utility. 👇 Read more below
Yuchen Jin@Yuchenj_UWAI 评分66The Decoder:AI News(RSS)精选AI 评分8282 Anthropic 发布 Claude Sonnet 5.5,基准接近 Opus 5.5 且单任务成本最多低 30%
Anthropic 发布 Claude 5.5 家族第二款模型 Claude Sonnet 5.5,输出速度提升超过 30%,单任务成本最多降低 30%,在多项基准上接近 Opus 5.5。
推荐理由:对比 Sonnet 5.5 与 Opus 5.5、Sonnet 5 的基准与定价,可看清中端模型在编码任务上的跃升与成本取舍。
Artificial Analysis@ArtificialAnlysAI 评分5656
Artificial Analysis@ArtificialAnlysAI 评分4747
ginobefun@hongming731AI 评分4040
dex@dexhorthyAI 评分55看腻了 slopcodebench 那些框框的进度,所以做了个……更潦草的东西

ginobefun@hongming731AI 评分4646
SemiAnalysis 长文 RSS(RSS)AI 评分5555 SemiAnalysis 拆解 Intel Panther Lake:18A 的 PowerVia、RibbonFET 与 Foveros-S 实现
SemiAnalysis 拆解 Intel Panther Lake,解析 Intel 18A 首个商用背面供电(PowerVia)、RibbonFET 四层纳米片和 Foveros-S 封装。实测显示 18A 计算逻辑密度与 TSMC N3E GPU 逻辑相近,但不及 N3P、N2 和 Samsung SF2 的峰值密度;NPU 5 面积缩小 36.9%,GPU 高端版本仍用 TSMC N3E。
karminski-牙医@karminski3AI 评分4343
elsewhere:文章(RSS)AI 评分5959 十字路口实测 Today AI:一周体验 Personal AI 助理
十字路口团队分享了对 Personal AI 产品 Today AI 为期一周的实测,覆盖 macOS、Windows、Linux、iOS 四平台。
SemiAnalysis 长文 RSS(RSS)AI 评分5959 SemiAnalysis 发布 ClusterMAX 3.0 GPU 云评级,覆盖 77 家 neocloud
SemiAnalysis 发布 ClusterMAX 3.0,覆盖 77 家 neocloud 供应商、323 家市场全景,并访谈超 200 名终端用户。Nebius 加入 CoreWeave 进入铂金级,Google Cloud 升入金级,Azure 降至银级,全球仅 19 家获得奖章评级,另新增介于铜级与不合格之间的 Participation Ribbon 级。
Epoch AI@EpochAIResearchAI 评分3939AI 能看出你把宜家家具装错了吗?我们的新基准——家具组装基准(FAB)——给模型提供说明书和一张半成品家具的照片,让它们找出错误。 最高分在短短 10 个月内从 28% 提升到了 80%。
OpenAI:官网动态(RSS · 排除企业/客户案例)AI 评分5656 OpenAI 发布开放基准 MentalHealthBench,评估 AI 心理健康对话能力
OpenAI 于 2026 年 9 月 23 日推出开放基准 MentalHealthBench,与来自 22 个国家、超过 80 名持证心理健康专家共同创建,用于评估 AI 在真实心理健康对话中的表现。
karminski-牙医@karminski3AI 评分3131
Perplexity@perplexity_aiAI 评分4545
Boris Cherny@bcherny精选AI 评分7070引用Claude@claudeaiIntroducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 for most tasks, and costs 40% less to run than Opus 5.
推荐理由:作者实测对比 Opus 5.5 与 Fable 5.1 移植 HAProxy 的用时和成本,给出第一手数据供选型参考。
karminski-牙医@karminski3AI 评分5454
elsewhere:文章(RSS)AI 评分6868 阶跃 Step 5 Preview 实测评测:数据可视化与金融分析亮眼,泛化和审美仍有短板
阶跃发布 Step 5 Preview,总参数量 600B、激活参数 27B,有视觉输入,官方称在 Artificial Analysis 上涨 44 分、单任务成本仅为 Claude Opus 5 的 1/8。作者与友人实测发现其在数据可视化、金融分析上表现不错,但泛化性、领域知识和审美偏弱,思考过程过长导致长任务耗时且易中断,且长上下文下安全指令遵循可被绕过。
Hugging Face:Blog(RSS)AI 评分3333 英国 AISI 如何用 EvalEval 基础设施让评测结果可复现
英国 AI Security Institute(AISI)正采用 EvalEval 的 Every Eval Ever schema 和 Evaluation Cards 平台公开评测结果。
eric zakariasson@ericzakariassonAI 评分3030引用aditya@adxtyahqGROK 4.7 IS ACTUALLY COMPETING WITH GPT-6 ASTRA. I gave GPT-6 Astra, Grok 4.7, Kimi K3 and Fable 5.1 the same prompt to build a flight simulator Astra was still #1 overall, but Grok 4.7 was surprisingly close Kimi K3 and Fable 5.1 were basically a draw and both produced a much smoother result, while Grok 4.7 was right up there with Astra in terms of overall quality. overall: Astra > Grok ≈ Kimi ≈ Fable GPT-6 Astra finally has some serious competition.
Logan Kilpatrick@OfficialLoganKAI 评分2323如果你在用 AI 构建产品,你应该花超过 25% 的时间来做基准测试,并努力让模型实验室关注这些基准测试 这是加速公司进展的最简单路径
Xiaomi MiMo@XiaomiMiMoAI 评分4040引用Design Arena@DesignArenaBREAKING: MiMo-V2.6-Pro by @XiaomiMiMo lands at #8 overall (#3 open-weight) on Design Arena with an Elo of 1338. This is an impressive 54-point and 22-position increase from MiMo-V2.5-Pro. MiMo-V2.6-Pro also reaches #4 overall in Website (#2 open-weight) and #6 overall in Agentic Frontend Development (#2 open-weight), showing strong performance across both direct generation and agentic coding. This places @XiaomiMiMo's new model among leading models such as GPT-5.6 Sol by @OpenAI, Claude Opus 5 by @Anthropic, and Kimi K3 by @MoonshotAI. Congratulations to the @XiaomiMiMo team on returning to a top-10 placement on Design Arena!
Epoch AI@EpochAIResearchAI 评分4949
Karina@karinanguyenAI 评分5656引用Epoch AI@EpochAIResearchIntroducing Benchmark Reviews: our new initiative to audit AI benchmarks. We are launching with 15 benchmarks: 4 Verified, 9 Flawed, and 2 with not enough information for a review.
Epoch AI@EpochAIResearchAI 评分5353
Epoch AI@EpochAIResearchAI 评分3232
Dan Hendrycks@hendrycksAI 评分4747
SemiAnalysis 长文 RSS(RSS)AI 评分5757 SemiAnalysis 实测 Vera Rubin NVL72 智能体推理:每美元 TCO 吞吐最高约 67 倍于 GB300
SemiAnalysis 在自建 AgentX 智能体推理基准上发布 Rubin 平台首批经核验的实测结果,称在 170 TPS、自有 TCO 口径下,Vera Rubin NVL72 相比 GB300 Dynamo TRTLLM 实现约 67 倍的总吞吐;在更常见的 60-100 TPS 区间为 1.4-3 倍。
Lee Robinson@leerobAI 评分2727我们刚刚推出了 CursorBench 4.0! 它包含了新的任务,用于评估模型遵循指令的能力、在具有挑战性的项目上长期工作的表现,并且比之前更难(所以所有模型的得分都更低了)。
jietang@jietangAI 评分1616SemiAnalysis 长文 RSS(RSS)AI 评分6565 SemiAnalysis 发布 InferenceX 预览:TPUv7 Ironwood 对比 Blackwell 每美元性能最高领先 50%
SemiAnalysis 在 InferenceX 官方预览中发布首个 TPUv7 Ironwood 第三方推理结果,FP8 聚合服务下每美元性能最高比 B200/B300 好 50%,20 tok/s/user 时每百万 token 成本 0.181 美元,低于 B200 的 0.222 美元和 B300 的 0.276 美元。
上海人工智能实验室 InternLM:原创项目AI 评分2929 上海人工智能实验室 InternLM 发布 SciDocBench 科学文档理解基准
上海人工智能实验室 InternLM 开源 SciDocBench,一个以工作流为中心的科学文档理解基准。该基准的官方仓库已公开,用于评测科学文档理解任务。
Logan Kilpatrick@OfficialLoganKAI 评分1919