Tencent Hy· @TencentHunyuan · X·· 2 小时前AI 评分56
AI 导读
腾讯混元、复旦大学与清华大学研究人员发布 ExplorationBench,一个测量 AI 系统探索能力的基准,论文见 https://arxiv.org/abs/2609.30199。基准构建规则可执行、与已知知识冲突的可验证外星世界,含 AlienCode(31 处隐藏规则变化、70 任务)和 AlienLogic(24 条修改的推理规则、70 定理)两个沙盒,答案由解释器或证明检查器判定,不用 LLM judge。对 10 个前沿 AI 系统的测试发现:有反馈时最佳系统四轮后从 15.7% 以下升至 89.0%,无反馈仅 0.5–11.0%;重放系统自己最佳探针时 10 个系统中 9 个表现更差;规则全部正确陈述时任务解决率仅 73.4%;同一系统同一预算下得分在 5.7% 到 79.0% 之间波动,两个世界间排名几乎不可迁移。代码即将开源:https://github.com/Tencent-Hunyuan/ExplorationBench
正文
New Research: We are releasing ExplorationBench, a benchmark for measuring how AI systems explore. Scientific discovery begins where known problems end: a system has to frame hypotheses, design experiments, and learn from the results. Evaluating this is hard. Genuinely new answers cannot be checked quickly, and in familiar domains a model can simply recall what it has seen. Addressing this challenge, researchers from Tencent Hy, Fudan University, and Tsinghua University built verifiable Alien Worlds. Their rules are executable, so every answer is checked exactly, and they conflict with familiar knowledge, so recall alone cannot solve the tasks. 🔹 Two sandboxes: AlienCode (31 hidden rule changes, 70 tasks) and AlienLogic (24 patched inference rules, 70 theorems) 🔹 A flawed manual, four rounds of self-designed probes, and closed-book tests after every round 🔹 Every answer graded by an interpreter or a proof checker, with no LLM judge What we found across 10 frontier AI systems: 1️⃣ Getting feedback is more effective than thinking alone. No AlienCode run starts above 15.7%; after four rounds the best reaches 89.0%, while the same turns without feedback stay at 0.5–11.0%. 2️⃣ Designing the experiments matters. Replaying a system's own best probes gives it exactly the same evidence, yet in AlienCode 9 of 10 systems do worse than when they chose the probes themselves. 3️⃣ Knowing a rule is not using it. Even when every required rule is stated correctly, tasks are solved only 73.4% of the time. 4️⃣ One score hides a lot. The same system under the same budget ended anywhere from 5.7% to 79.0%, and rankings barely transfer between the two worlds. CL-bench asked whether models can learn from context. ExplorationBench asks whether they can discover the rules themselves. 📄 Paper: https://arxiv.org/abs/2609.30199 🌐 Website & leaderboard: https://explorationbench.com 📝 Blog: https://explorationbench.com/blog/ 💻 Code (coming soon): https://github.com/Tencent-Hunyuan/ExplorationBench在 X 查看被引用的帖子
来源:Tencent Hy · x.com