Arena 展示了从 2025 年 Q4 至今前沿模型在"复现古罗马"任务上的进展,Claude Sonnet 5.5 与 GPT-6.1 的分数即将公布。参与评测的模型包括 Claude Opus 4.5 至 5.5、Claude Fable 5/5.1、GPT-5.2 至 6 Sol、Gemini 3.1、Kimi K3 和 Qwen 3.8 Max,榜单任务来自其全球用户社区的真实需求。
Watch the progress of frontier models in bringing Ancient Rome to life on Arena, from Q4 2025 to now.
Scores for Claude Sonnet 5.5 by @AnthropicAI and GPT-6.1 by @OpenAI are coming soon. Real-world tasks from our global community of users power the Arena leaderboards. Head to Arena now to test it out, and stay tuned!
Featured models:
- Claude Opus 4.5, 4.6, 4.7, 4.8, 5 and 5.5
- Claude Fable 5 and 5.1
- Claude Sonnet 5.5
- GPT-5.2, 5.3 Codex, 5.4, 5.5, 5.6 Sol, 6 Astra, and 6 Sol
- Gemini 3.1
- Kimi K3
- Qwen 3.8 Max
Find the prompt from @petergostev below.
来源:Arena.ai · x.com