arXiv:cs.LG(机器学习,全量分类)· Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")·· 1 天前AI 评分34
无需架构创新:大规模遥感 VLM 用简单配方实现 SOTA
More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe
AI 导读
一项被 ACCV 2026 接收的研究提出,通用视觉语言模型只要在多样数据与任务上以足够规模训练,无需遥感专用架构即可在多个遥感基准上达到竞争性或 SOTA 表现。
正文
Abstract:Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks. Most recent progress in this area has been driven by remote-sensing-specific architectural designs, often introducing new encoders, alignment modules, or task-specific fusion mechanisms. In this work, we challenge the necessity of such architectural specialization. We show that a generally capable vision-language model can achieve competitive or state-of-the-art performance at challenging remote sensing benchmarks, provided that it is trained at sufficient scale across diverse data and tasks. Our model uses a single language policy that can either answer directly in text or invoke a localization tool for segmentation and grounding. To train this heterogeneous behaviour, we employ a multi-task reinforcement learning framework with adaptive task rewards covering multiple-choice VQA, free-form VQA, captioning, detection, and segmentation across a large variety of input types. Our approach achieves competitive results across a broad set of benchmarks, including high-resolution, multi-temporal, multi-modal and multi-view tasks. Further, as training data scales, our experiments show consistent improvements across most tasks both in and out of distribution, which correlate with per-task data diversity. These findings suggest that, for remote sensing VLMs, data scale is sufficient even without architectural novelty.
| Comments: | ACCV 2026. Project Page this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.15942 [cs.CV] |
| (or arXiv:2607.15942v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15942 arXiv-issued DOI via DataCite |
Submission history
From: Stefan Maria Ailuro [view email]
[v1]
Fri, 17 Jul 2026 13:25:44 UTC (12,414 KB)
[v2]
Thu, 1 Oct 2026 09:51:10 UTC (12,423 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org