arXiv:cs.LG· Wenhao Li, Xianjing Meng, Qiangchang Wang, Zhongyi Han, Yilong Yin, Liqiang Nie·· 4 小时前
DVLA-RL++:用强化学习门控实现双层视觉-语言对齐的小样本学习
DVLA-RL++: Dual-Level Vision-Language Alignment with Reinforcement Learning Gating for Few-Shot Learning
AI 导读
DVLA-RL++ 在 DVLA-RL 基础上加入互补语义净化(CSP)与反事实强化学习门控(CRG),通过区分内在描述与干扰描述并按歧义程度稀疏分配证据、以内在语义锚点兜底,同时用平衡识别性能与干扰暴露的奖励学习逐层语义融合强度。在标准、细粒度和跨域 benchmark 上取得 SOTA 准确率,平均较 DVLA-RL 提升 1.4%。
正文
Abstract:Few-shot learning aims to recognize novel categories from limited labeled examples. Recent studies incorporate textual semantics to compensate for limited visual observations and improve class representations. However, high image-text agreement may reflect both intrinsic object properties and incidental context, making support prototypes susceptible to contextual contamination. To address this problem, we propose DVLA-RL++, which extends DVLA-RL with complementary semantic purification (CSP) and counterfactual reinforcement-learning gating (CRG). Specifically, CSP generates intrinsic and nuisance descriptions from labeled supports and compares their agreement with each support token. An ambiguity-dependent rejection margin guides sparse evidence allocation, while an intrinsic semantic anchor fills the unassigned mass to provide a fallback when visual evidence is unreliable. CRG learns layer-wise semantic fusion strengths using a reward that balances recognition performance and nuisance exposure. An independently executed reference trajectory on the same episode provides a paired learning signal. Theoretical analysis relates retained evidence and anchor quality to prototype stability and establishes conditions for unbiased on-policy gradient estimation. Experiments on standard, fine-grained, and cross-domain benchmarks show state-of-the-art accuracy, with an average gain of 1.4% over DVLA-RL. The project page is available at this https URL.
| Comments: | This work has been submitted to the IEEE TPAMI for possible publication |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.12095 [cs.CV] |
| (or arXiv:2610.12095v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.12095 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Wenhao Li [view email]
[v1]
Thu, 8 Oct 2026 15:01:47 UTC (3,794 KB)
来源:arXiv:cs.LG · arxiv.org