Arena 宣布 OpenAI 的 GPT-6.1 已上线,用户可前往 Agent Arena 测试,投票将影响其评估,分数即将公布。Agent Arena 基于数百万真实世界长程智能体任务测量模型,模型可使用 web 搜索、文件系统和终端工具完成复杂工作流,排行榜用 causal tracing 方法衡量模型相对平均模型的结果表现。
Arena 官宣 GPT-6.1 上线 Agent Arena 和 Code Arena,读者可实际参与投票并等待基于真实智能体任务的评测分数。
GPT-6.1 by @OpenAI is now in the Arena!
Head to Agent Arena to test it out, and your votes will shape its evaluation. Scores coming soon.
In Agent Arena, we measure models on millions of real-world, long-horizon agentic tasks. Models can access web search, filesystem, and terminal tools to complete complex workflows. The leaderboard measures model performance on outcomes relative to the average model using a causal tracing methodology.
GPT-6.1 is also available in Code Arena: WebDev, Text, Vision, Search and Document.
GPT-6.1 Sol: near-Astra intelligence for a fifth of the price. It’s the most cost-efficient model for its performance available today.在 X 查看被引用的帖子
来源:Arena.ai · x.com