跳到正文
arXiv:cs.AI· Uttamasha Monjoree, Wei Yan·· 10 小时前AI 评分29

Agentic AI 结合结构化 CoT 提升空间智能:3D 旋转的可视化与推理

Agentic AI with Structured CoT for Enhancing AI's Spatial Intelligence: Visualization and Reasoning of Rotation

AI 导读

研究用改良版 Revised PSVT:R 测试 GPT-5.6 的空间旋转推理能力,发现结构化 CoT 在 PSVT:R 及带坐标系两个数据集上均提升基础模型表现。

正文

View PDF HTML (experimental)

Abstract:Recent studies show that artificial intelligence (AI) with language and vision capabilities still experiences limitations in spatial reasoning. In this paper, we have studied the spatial capabilities of advanced generative AI to understand the rotations of objects in 3D space, utilizing AI's image processing and language processing features. We trained and examined the spatial intelligence of a generative Agentic AI model (GPT-5.6) to understand the spatial rotation process with rotation diagrams based on the revised Purdue Spatial Visualization Test: Visualization of Rotations (Revised PSVT:R). We improvised the Revised PSVT:R by superimposing additional graphical and contextual features to evaluate how different Chain-of-Thought (CoT) reasoning strategies influence model performance. The results indicate that structured CoT reasoning improves the spatial reasoning performance of the base GPT-5.6 model in both datasets (PSVT:R and PSVT:R with coordinate system). We used three CoT approaches - (1) Structured CoT, (2) few-shot Structured CoT, and Structured CoT with Self-optimized Prompt. The three CoT approaches evaluated in this study showed no significant performance difference. Results showed that combining structured CoT reasoning with relevant contextual information leads to considerable improvements in VLM performance on 3D rotation tasks, demonstrating the potential of agentic AI for more effective spatial reasoning. However, when contextual information is removed, structured CoT reasoning alone provides limited improvement, and the models continue to exhibit notable difficulties in understanding spatial transformations. These findings suggest that effective spatial reasoning in VLMs relies on the integration of visual, textual, and reasoning-based information in future agentic AI systems for spatial intelligence.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.04188 [cs.AI]
  (or arXiv:2610.04188v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.04188

arXiv-issued DOI via DataCite

Submission history

From: Uttamasha Monjoree [view email]
[v1] Sat, 3 Oct 2026 00:56:55 UTC (1,875 KB)
[v2] Tue, 6 Oct 2026 03:11:55 UTC (1,875 KB)

来源:arXiv:cs.AI · arxiv.org