跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Yaodong Yang, Hongyao Tang, Yi Ma, Xingyu Fan, Weixun Wang, Jinpeng Li, Tianpei Yang·· 15 小时前AI 评分58

arXiv 论文:Jev 判断模型擅长评估但不擅长模拟,代码补上模拟后成为专家控制器

Code Owns the Simulation, Jev Owns the Evaluation

AI 导读

arXiv 论文(arXiv:2610.01834)测试判断模型 Jev 在反思测试、矩阵博弈、ALFWorld 和机器人控制上的表现,发现清晰边界:当正确选项可从输入描述中判断时(评估)。

正文

View PDF HTML (experimental)

Abstract:Judgment models such as \jev{} return, in a single call and without reasoning text, a probability for each described option. This makes them attractive as an agent's action-selection layer, but it is unclear which decisions they can be trusted with. We test \jev{} on reflection tests, one-shot matrix games, the text game ALFWorld and robot control, and find a sharp boundary. \jev{} succeeds when the right option can be judged from what the input describes, which we call \emph{evaluation}. Specifically, it solves 99\% of the counterintuitive Cognitive Reflection Test questions. However, it fails when the right option depends on \emph{simulation} (i.e., predicting something not in the input), such as the opponent's action or the subgoal that must come first. In games, \jev{} plays suboptimally as if its rational opponent acted at random, because the opponent's action is not given. In ALFWorld, \jev{} favors commands that mention an object or place named in the task description. For example, given the task ``put a clean knife in the drawer'', \jev{} carries an unwashed knife straight to the drawer instead of first washing it at the sink. Surprisingly, many of these failures are not due to a lack of knowledge. Asked separately what the opponent will do, \jev{} usually answers correctly, and it responds well given the opponent's action. It fails when one call must both perform the simulation and evaluate based on it. This suggests letting code make the prediction or simulation. When code supplies it, such as a lookahead in ALFWorld and physics simulation in robot control, \jev{} becomes an expert controller through its general evaluation ability.
Comments: 10 pages main text, 20 pages total with appendix; 6 figures, 7 tables. Preprint
Subjects: Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2610.01834 [cs.AI]
  (or arXiv:2610.01834v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.01834

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yaodong Yang Mr. [view email]
[v1] Thu, 1 Oct 2026 15:06:40 UTC (411 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org