arXiv:cs.LG· Johannes K\"unzel, Peter Eisert, Anna Hilsmann·· 4 小时前AI 评分38
RIPE++:仅用正样本对的强化学习关键点学习方法
RIPE++: Reinforced Keypoint Learning from Positive Pairs Only
AI 导读
针对 RIPE 依赖粗粒度二值奖励和负样本对的问题,RIPE++ 提出从单个正样本对同时导出奖励与惩罚,无需负样本对比即可学习判别性检测器和描述子。该方法将该 RL 目标扩展到匹配阶段并适配 LightGlue,使 MegaDepth1500 上的 AUC@5 从 56.58 提升至 59.65,并支持仅用部分视觉重叠图像对进行弱监督训练。方法在低纹理医学视频序列上同样可训练,代码与数据已公开。
正文
Abstract:Sparse keypoint extraction and matching underpin core tasks in geometric computer vision, including structure-from-motion, visual SLAM, augmented reality, and medical image registration. Learning robust local feature representations, however, typically requires accurate camera poses or depth supervision, which are often unavailable in real-world settings. Reinforcement learning (RL) has recently emerged as a promising alternative, requiring only the information if two images show the same scene or not. However, existing RL formulations such as RIPE rely on coarse binary rewards and carefully constructed negative training pairs, limiting training stability and descriptor discriminability. In this paper, we revisit RL-based keypoint learning and propose a reward that fully exploits the geometric consistency signal, deriving both reward and penalty from a single positive pair without contrasting against negatives. This richer signal provides sufficient supervisory contrast to learn discriminative detectors and descriptors from positive image pairs alone, enabling representation learning under extremely limited supervision. Furthermore, we show that the same RL objective can be extended to the matching stage by adapting LightGlue, raising AUC@5 on MegaDepth1500 from 56.58 to 59.65 and enabling weakly-supervised training of the full sparse matching pipeline from image pairs with partial visual overlap. We validate our approach on established benchmarks, demonstrating competitive results compared to fully-supervised methods. We further show that the method can be even trained on low texture medical video sequences, where camera poses are usually unavailable and standard SfM pipelines often fail. Code and data are available at this https URL .
| Comments: | LIMIT@ECCV 2026 (Best Paper Award) |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2608.19693 [cs.CV] |
| (or arXiv:2608.19693v2 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.19693 arXiv-issued DOI via DataCite |
Submission history
From: Johannes Wolf Künzel [view email]
[v1]
Thu, 20 Aug 2026 06:37:24 UTC (35,584 KB)
[v2]
Tue, 6 Oct 2026 09:29:53 UTC (35,584 KB)
来源:arXiv:cs.LG · arxiv.org