跳到正文
原文
Arena.ai· @arena · X·· 3 小时前精选AI 评分73
AI 导读

Arena 宣布 Claude Sonnet 5.5 (xHigh) 进入 Code Arena: WebDev 榜单,以 1786 分排名第 3,距第 2 名 GPT-6 Astra(1788 分)仅 2 分。

推荐理由

榜单数据给出 Claude Sonnet 5.5 (xHigh) 的排名与价格对比,读者可据此权衡性价比和选型。

正文 · AI 翻译

激动人心的更新:具备 xHigh 推理能力的 Claude Sonnet 5.5 已登陆 Code Arena: WebDev。以 1786 分位列第 3!

在每百万 token 综合成本 8 美元的情况下,Claude Sonnet 5.5 凭借 xHigh 推理能力依然处于帕累托前沿。此次发布距离排名第 2、得分 1788 的 GPT-6 Astra 仅差 2 分,而价格只有其 80%。

按领域划分,Claude Sonnet 5.5(xHigh)的成绩为:
- 在 Gaming、Reference-Based Design 和 Brand & Marketing 中排名第 2
- 在 Simulations 中排名第 3
- 在 Content Creation Tools 和 Consumer Product 中排名第 4
- 在 Data & Analytics 中排名第 6

再次祝贺 @AnthropicAI 发布此次更新!

引用Arena.ai@arena
@AnthropicAI 的 Claude Sonnet 5.5 (High) 真实世界结果出炉了。它刚在 Code Arena: WebDev 中以 1699 分位列第 4,并凭借其成本效率重塑了帕累托前沿! Claude Sonnet 5.5 (High) 以每百万 token 8 美元的混合成本实现了近乎顶级的性能,重塑了帕累托前沿!该模型比总排名第 3 的 Claude Fable 5.1 (Max) 和排名第 2 的 GPT-6 Astra (Max) 都便宜 80%。 见下方帕累托位置。 总体而言,Claude Sonnet 5.5 (High) 相比排名第 37、得分 1540 的 Sonnet 5 (High) 提升了 +159 分。与其前代变体相比的这一提升,目前也体现在以下关键领域: - 基于参考的设计:#38 → #4 - 模拟:#37 → #4 - 游戏:#36 → #4 祝贺 @AnthropicAI 发布!
原文

Real-world results are in for Claude Sonnet 5.5 (High) by @AnthropicAI. It just landed #4 in Code Arena: WebDev with 1699 pts, and has reshaped the Pareto frontier with its cost efficiency! Claude Sonnet 5.5 (High) delivers nearly top performance at a blended $8 per Mtoken, reshaping the Pareto frontier! This model is 80% cheaper than both Claude Fable 5.1 (Max) in the #3 spot overall, and GPT-6 Astra (Max) at #2. See Pareto placement below. Overall, Claude Sonnet 5.5 (High) is a +159 pt improvement from Sonnet 5 (High) at #37 with 1540 pts. This gain compared to its previous variant also shows up across these key domains so far: - Reference-Based Design: #38 → #4 - Simulations: #37 → #4 - Gaming: #36 → #4 Congrats to @AnthropicAI on this release!

在 X 查看被引用的帖子

来源:Arena.ai · x.com