arXiv:cs.AI· Uttamasha Monjoree, Wei Yan·· 10 小时前AI 评分32
微调 VLM 提升 AI 空间智能:理解 3D 与 2D 旋转
Fine-Tuning VLM for Enhancing AI's Spatial Intelligence: Understanding 3D and 2D Rotations
AI 导读
研究通过多组物体旋转数据集微调 Google DeepMind 的 Gemma-4 MoE 模型,在 2D 和 3D 旋转检测上均取得提升。微调后的 Gemma-4 MoE 在按轴和角度预测旋转方面显著优于微调后的 Gemma-4 通用模型,且无需显式坐标系即可改善 2D 角度估计。实验还发现,可识别物体并未提升角度检测准确率,具有显著线性特征的物体反而表现更好。
正文
Abstract:Spatial intelligence is a fundamental skill in multiple domains, such as Science, Technology, Engineering, and Mathematics (STEM), Medicine, Architecture, and Construction. Recent studies indicate that Vision-Language Models (VLMs) still face limitations in spatial reasoning, which inhibits artificial intelligence (AI) from performing practical spatial tasks. Using multiple object-rotation datasets developed for training and evaluation, our experiments demonstrated promising improvements in both 2D and 3D rotation detection. Fine-tuned Google DeepMind-built Gemma-4 mixture-of-experts (MoE) models significantly outperformed fine-tuned Gemma-4 generalist models in predicting rotations defined by both their axes and angles. Fine-tuning also substantially improved angle estimation for 2D representation without requiring an explicit coordinate system. Furthermore, identifiable objects did not improve angle-detection accuracy; instead, objects with prominent linear features showed improved performance.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.04206 [cs.AI] |
| (or arXiv:2610.04206v2 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.04206 arXiv-issued DOI via DataCite |
Submission history
From: Uttamasha Monjoree [view email]
[v1]
Sat, 3 Oct 2026 01:44:27 UTC (1,497 KB)
[v2]
Tue, 6 Oct 2026 02:46:30 UTC (936 KB)
来源:arXiv:cs.AI · arxiv.org