arXiv:cs.LG· Ayesh Abu Lehyeh, Jay Hwasung Jung, Safwan Wshah·· 4 小时前AI 评分34
语言能为跨视角地理定位保留多少信息?MLLM 零样本推理研究
What Words Keep of a Place: Zero-Shot Language Reasoning for Cross-View Geo-Localization
AI 导读
研究用多模态大语言模型(MLLM)将地面全景与卫星瓦片描述为结构化文本,以语言比较完成跨视角地理定位,全程无需训练。在 VIGOR 的 9,826 对样本上,全库按描述相似度排序的 Recall@1 仅 0.39%;将候选缩小到十块相邻瓦片后,MLLM 评判结构一致性使排名效果翻倍并追平强词法基线,还能指出哪些字段一致或冲突。
正文
Abstract:Cross-view geo-localization is commonly solved as an image retrieval problem, matching a ground-level image against a database of satellite tiles through a jointly trained embedding. Such models are accurate, but they need large paired supervision and cannot show what evidence supports a match. In this paper, we study a different question: how much of this task can be solved through language alone? We prompt a multimodal large language model (MLLM) to describe each ground panorama and each satellite tile as structured text, and localize by comparing these descriptions. No component is trained. We evaluate on 9,826 VIGOR pairs from four U.S. cities, in three settings. First, the descriptions are faithful but not discriminative. They agree closely across the two views, yet ranking the full pool by description similarity almost never returns the correct tile (0.39% Recall@1). Second, we narrow the pool to ten neighboring tiles, as a coarse prior would do. The same descriptions now become useful: an MLLM judge that scores structural consistency doubles random ranking and matches a strong lexical baseline. It also states which fields of the two descriptions agree and which conflict, which an embedding distance cannot do, and which we see as a step toward interpretable localization. Third, we place the judge on a trained visual retriever. On the queries it ranks wrongly, reranking from images works, while reranking from our descriptions does not (23.5% against 10.7% Recall@1). Scene structure survives the conversion into language, while the fine appearance detail needed to separate nearby places does not. Code and prompts are publicly available at this https URL.
| Comments: | Accepted at NeurIPS 2026 Workshop Physical World AI: Geometry, Characteristics, and Multimodal Sensing |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.07269 [cs.CV] |
| (or arXiv:2610.07269v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07269 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ayesh Abu Lehyeh [view email]
[v1]
Mon, 5 Oct 2026 19:11:38 UTC (2,522 KB)
来源:arXiv:cs.LG · arxiv.org