跳到正文
arXiv:cs.LG· Yunsong Fang, Tingyu Wang, Zhedong Zheng·· 3 小时前

GeoFuse:利用道路地图作为免费几何先验的天气不变无人机地理定位

Road Maps as Free Geometric Priors: Weather-Invariant Drone Geo-Localization with GeoFuse

AI 导读

GeoFuse 是一个跨模态融合框架,将精确对齐的道路地图瓦片与卫星影像结合,以提升恶劣天气下无人机图像地理定位的鲁棒性。作者为 University-1652 和 DenseUAV 基准补充了地理对齐道路地图,并设计融合模块通过 token 级和通道级交互与动态门控机制自适应加权各模态贡献。实验显示其在两个基准上的 Recall@1 分别提升 3.46% 和 23.18%。

正文

View PDF HTML (experimental)

Abstract:Drone-view geo-localization aims to match a query drone image, often captured under adverse weather conditions (e.g., rain, snow, fog), against a gallery of geo-tagged satellite images. Weather-induced degradations in the drone view, such as noise, reduced visibility, and partial occlusions, severely exacerbate the intrinsic cross-view domain gap. While prior methods predominantly rely on weather-specific architectures or data augmentations, they have largely overlooked road map data, a readily available modality that provides strong, inherently weather-invariant geometric layout cues (e.g., road networks and building footprints) at negligible additional cost. We introduce GeoFuse, a cross-modal fusion framework that integrates precisely aligned road map tiles with satellite imagery to yield more discriminative and weather-resilient representations. We first augment the existing University-1652 and DenseUAV benchmarks with geo-aligned road maps, supplying structural priors robust to meteorological variations. Building on this, we propose a flexible fusion module that combines satellite and road map features via token-level and channel-level interactions, with a lightweight dynamic gating mechanism that adaptively weights modality contributions per instance. Finally, we employ class-level cross-view contrastive learning to promote robust alignment between weather-degraded drone features and the fused satellite-roadmap representations. Extensive experiments under diverse weather conditions show that GeoFuse consistently outperforms state-of-the-art methods, achieving +3.46% and +23.18% Recall@1 accuracy on the University-1652 and DenseUAV benchmarks, respectively.
Comments: 18 pages, 4 figures
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2605.14925 [cs.CV]
  (or arXiv:2605.14925v3 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2605.14925

arXiv-issued DOI via DataCite

Submission history

From: Yunsong Fang [view email]
[v1] Thu, 14 May 2026 15:01:22 UTC (1,598 KB)
[v2] Mon, 29 Jun 2026 02:02:31 UTC (1,598 KB)
[v3] Thu, 8 Oct 2026 16:01:30 UTC (1,598 KB)

来源:arXiv:cs.LG · arxiv.org