跳到正文
原文
arXiv:cs.LG(机器学习,全量分类)· Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")·· 1 天前AI 评分34

无需架构创新:大规模遥感 VLM 用简单配方实现 SOTA

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

AI 导读

一项被 ACCV 2026 接收的研究提出,通用视觉语言模型只要在多样数据与任务上以足够规模训练,无需遥感专用架构即可在多个遥感基准上达到竞争性或 SOTA 表现。

正文

View PDF HTML (experimental)

Abstract:Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at sufficient scale across diverse data and tasks. Our model uses a single language policy that can either answer directly in text or invoke a localization tool for segmentation and grounding. To train this heterogeneous behaviour, we employ a multi-task reinforcement learning framework with adaptive task rewards covering multiple-choice VQA, free-form VQA, captioning, detection, and segmentation across a large variety of input types. Our approach achieves competitive results across a broad set of benchmarks, including high-resolution, multi-temporal, multi-modal and multi-view tasks. Further, as training data scales, our experiments show consistent improvements across most tasks both in and out of distribution, which correlate with per-task data diversity. These findings suggest that, for remote sensing VLMs, data scale is sufficient even without architectural novelty.
Comments: ACCV 2026. Project Page this https URL
Subjects: Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG)
Cite as: arXiv:2607.15942 [cs.CV]
  (or arXiv:2607.15942v2 [cs.CV] for this version)
  https://doi.org/10.48550/arXiv.2607.15942

arXiv-issued DOI via DataCite

Submission history

From: Stefan Maria Ailuro [view email]
[v1] Fri, 17 Jul 2026 13:25:44 UTC (12,414 KB)
[v2] Thu, 1 Oct 2026 09:51:10 UTC (12,423 KB)

来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org