跳到正文
arXiv:cs.AI· Uttamasha Monjoree, Wei Yan·· 10 小时前AI 评分32

微调 VLM 提升 AI 空间智能:理解 3D 与 2D 旋转

Fine-Tuning VLM for Enhancing AI's Spatial Intelligence: Understanding 3D and 2D Rotations

AI 导读

研究通过多组物体旋转数据集微调 Google DeepMind 的 Gemma-4 MoE 模型,在 2D 和 3D 旋转检测上均取得提升。微调后的 Gemma-4 MoE 在按轴和角度预测旋转方面显著优于微调后的 Gemma-4 通用模型,且无需显式坐标系即可改善 2D 角度估计。实验还发现,可识别物体并未提升角度检测准确率,具有显著线性特征的物体反而表现更好。

正文

View PDF HTML (experimental)

Abstract:Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in spatial reasoning, which inhibits artificial intelligence (AI) from performing practical spatial tasks. Using multiple object-rotation datasets developed for training and evaluation, our experiments demonstrated promising improvements in both 2D and 3D rotation detection. Fine-tuned Google DeepMind-built Gemma-4 mixture-of-experts (MoE) models significantly outperformed fine-tuned Gemma-4 generalist models in predicting rotations defined by both their axes and angles. Fine-tuning also substantially improved angle estimation for 2D representation without requiring an explicit coordinate system. Furthermore, identifiable objects did not improve angle-detection accuracy; instead, objects with prominent linear features showed improved performance.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.04206 [cs.AI]
  (or arXiv:2610.04206v2 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.04206

arXiv-issued DOI via DataCite

Submission history

From: Uttamasha Monjoree [view email]
[v1] Sat, 3 Oct 2026 01:44:27 UTC (1,497 KB)
[v2] Tue, 6 Oct 2026 02:46:30 UTC (936 KB)

来源:arXiv:cs.AI · arxiv.org