跳到正文
arXiv:cs.AI· Sadig Gojayev, Carolina Fortuna·· 3 小时前

面向智能序贯决策的 3D 表征框架:对比 Neurosolver、FBRL 与 AutoToS

A 3D Characterization Framework for Intelligent Sequential Decision Making

AI 导读

研究人员提出一个三维表征框架,将 AI 序贯决策方法投射到 MDP 形式化、设计的人类先验自主度、技能与计算成本三个维度进行统一比较。基于汉诺塔基准,该框架对比了图搜索方法 Neurosolver、前向-后向强化学习(FBRL)与基于 LLM 的 AutoToS(含双智能体扩展 DA-ToS)。

正文

View PDF HTML (experimental)

Abstract:Puzzles are widely used to evaluate the reasoning capabilities of artificial intelligence (AI) systems for sequential decision making, yet approaches originating from different paradigms are rarely compared under unified conditions. To address this gap, we introduce a three-dimensional characterization framework that enables the analysts of AI methods by 1) projecting them to the Markov decision process (MDP) sequential decision making formalism, 2) degree of autonomy through human prior ranking of their designs and, 3) skill and computational cost. Using this framework, we analyze how representative graph-based, reinforcement learning, and large language model (LLM)-based approaches differ in their design choices and performance characteristics, instantiated respectively by Neurosolver, forward-backward reinforcement learning (FBRL), and automated thought-of-search (AutoToS), including a double-agent extension of thought-of-search (DA-ToS). The analysis relies on the Tower of Hanoi puzzle that provides a controlled benchmark with well-defined rules and scalable complexity, enabling consistent comparison across increasing problem sizes. The 3D characterization reveals that LLM-based methods, due to their weakly constrained action-space design, shift complexity from architecture to inference-time verification, leading to substantially higher memory and runtime costs than Neurosolver and FBRL.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11696 [cs.AI]
  (or arXiv:2610.11696v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11696

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sadig Gojayev [view email]
[v1] Thu, 8 Oct 2026 11:06:51 UTC (1,002 KB)

来源:arXiv:cs.AI · arxiv.org