跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Arun Sharma·· 17 小时前AI 评分30

Spatial Atlas:面向空间感知研究智能体基准的计算锚定推理

Spatial Atlas: Compute-Grounded Reasoning for Spatial-Aware Research Agent Benchmarks

AI 导读

Spatial Atlas 提出计算锚定推理(CGR)设计模式,让代码先基于显式中间表示计算部分子问题,再由语言模型作答,并以 Agent2Agent(A2A)服务器形式实现空间问答与机器学习工程两个处理器。

正文

View PDF HTML (experimental)

Abstract:We describe compute-grounded reasoning (CGR), a design pattern in which code computes selected sub-problems from explicit intermediate representations before a language model answers. Spatial Atlas implements CGR as an Agent2Agent (A2A) server with a spatial question-answering handler and a machine-learning engineering handler. The spatial handler asks a language model to extract a scene graph, and code then fills in missing distances and checks the extracted safety rules. A separate benchmark driver can also run a strict metric bridge. It computes the gap for horizontal-gap questions from segmentation masks and a reconstructed point map, and it passes that gap to the answering model as a fact. The bridge returns a fixed unavailable answer when an evidence check fails, and it never falls back to model-estimated coordinates. The ML-engineering handler generates pipeline code, parses validation scores, and caps the number of repair and refinement passes. Its code execution is off by default. The repository also provides four run modes that can write label-free journals, a shuffled-image control mapping, and journal validators that reject label-bearing fields. We report one private label-free operational run in which four paths each wrote eight prediction rows with zero retries. Labels stayed sealed, and no score was computed, so this run establishes operational integrity only. We report no FieldWorkArena result because the benchmark data were not accessible. We also omit every performance, latency, and resource-use number that lacks a reproducible run artifact.
Comments: 11 pages. Code: this https URL
Subjects: Artificial Intelligence (cs.AI); Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2604.12102 [cs.AI]
  (or arXiv:2604.12102v3 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2604.12102

arXiv-issued DOI via DataCite

Submission history

From: Arun Sharma [view email]
[v1] Mon, 13 Apr 2026 22:22:07 UTC (20 KB)
[v2] Wed, 15 Apr 2026 03:29:47 UTC (20 KB)
[v3] Wed, 30 Sep 2026 18:06:26 UTC (37 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org