Arena 对比了 12 个模型在 1,460 场对战中的 34,580 次评审裁决与人类投票,发现 AI 评审的偏好与人类不同。AI 评审平均 58% 的时间选自己的答案,而人类只选同一答案 34%;GPT-6 Astra 自选率高达 88%。
Ask an AI model to pick the better answer, and it'll usually pick its own.
We compared 34,580 verdicts from 12 models with human votes across 1,460 battles on Arena. The results show that AI judges have their own taste, and it's unlike ours.
They:
- Favor their own answers. On average, a model picked its own answer 58% of the time. People picked that same answer 34% of the time. GPT-6 Astra picked itself 88% of the time.
- Rarely call a draw. People called a tie or "both bad" in 32% of battles. GPT-5.6 Sol picked a winner 96% of the time.
- Side with each other over people. They agreed with other AIs 79% of the time and with people 57% of the time. Every judge did, by a margin of 18 to 27 percentage points.
Full results in the article from @DawidGalarowicz below.
https://x.com/i/article/2104955648664584192在 X 查看被引用的帖子
来源:Arena.ai · x.com