跳到正文
arXiv:cs.CL· Jingyao Zhang, Yun Li, Lu Han·· 3 小时前AI 评分31

GAGR-Lab:评估联合空间几何与解析函数推理

GAGR-Lab: Evaluating Joint Spatial-Geometric and Analytic Function Reasoning

AI 导读

GAGR-Lab 是一个通过笛卡尔游戏场景、显式函数语义和 Rust 轨迹执行来测量联合空间几何与解析函数推理能力的框架。对 Llama 3.2 11B Vision Instruct 的有限试点运行了 72 个平衡游戏、432 次尝试,无一命中目标。特权解析搜索对照在 300 个生成场景的 600 个方向案例上独立成功,完整难度矩阵与多模型对比结果仍待测试。

正文

View PDF HTML (experimental)

Abstract:Joint spatial-geometric and analytic function reasoning requires translating a perceived spatial configuration into a symbolic function whose executed curve satisfies geometric constraints. We present GAGR-Lab, a framework for measuring this capability through Cartesian game scenes, explicit function semantics, and authoritative Rust trajectory execution. It distinguishes spatial perception, metric grounding, geometric relations, function interpretation, function construction, and constrained synthesis. We specify four configurable scene-difficulty presets and a prospective 24-cell diagnostic design, while reporting only the subset actually evaluated. A bounded pilot of one hosted model (Llama 3.2 11B Vision Instruct) using two API credentials as execution replicas yields 72 balanced games with 432 attempts, 429 valid provider responses, and no target hits; exploratory ordinary-function prompt variants also fail to hit, while the structured localization interface yields no scoreable outputs. A privileged analytic search control independently succeeds on 600 directional cases from 300 generated scenes, with exact repeatability and 1,200 successful vertical-reflection or translation checks. The framework separates serving reliability, symbolic compliance, and geometric success, and preserves exact model-visible inputs and realized paths. A staged protocol outlines diagnostic calibration, held-out replication, multi-model comparison, and paired robustness tests. The contribution is an operational research framework with an executed pilot and a clearly identified prospective study plan; the full difficulty matrix and comparative model results remain untested.
Comments: 15 pages, 1 figure, 7 tables
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2610.10201 [cs.AI]
  (or arXiv:2610.10201v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.10201

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Yun Li [view email]
[v1] Wed, 7 Oct 2026 15:01:54 UTC (75 KB)

来源:arXiv:cs.CL · arxiv.org