arXiv:cs.LG· Bill Psomas, Mohammad Mahdi, Michalis Thomas, Danda Pani Paudel, Giorgos Tolias, Giorgos Kordopatis-Zilos·· 4 小时前AI 评分40
面向表示学习的全局平均精度 gAP 与可微替代损失 gSAP
Global Average Precision for Representation Learning
AI 导读
研究者提出 Global Average Precision(gAP),把全部查询-候选对放进一个列表计算单个 AP,以衡量相似度在不同查询间是否可比;同时给出可微替代损失 gSAP,仅需相似度矩阵和正样本对二值矩阵,可直接替换 InfoNCE 等现有损失。
正文
Abstract:Standard information retrieval metrics, such as mean Average Precision (mAP), assess performance one query at a time, based on how the similarities between a query and its positives compare against those with its negatives. The same holds for common representation learning losses, such as InfoNCE and per-query AP surrogates. None of them considers whether similarities are comparable across queries, which any system with a single decision threshold relies on. Global Average Precision (gAP) does, by ranking all query-candidate pairs in one list and computing a single AP. We introduce gSAP, a differentiable surrogate of gAP. It needs only a similarity matrix and a binary matrix marking the positive pairs, the same input as existing losses, so it is a drop-in replacement for them and agnostic to the encoder, the modality, and the source of supervision. Since it considers all possible pairwise comparisons in the batch jointly, it also remains trainable at low temperatures, a regime where per-query surrogates run out of gradient. Swapping it into established recipes improves supervised metric learning, cross-modal alignment, and self-supervised pretraining, where, to our knowledge, it is the first ranking loss to replace the community standard InfoNCE in the latter two. Its similarities are more consistent across queries, which drives the gains under a universal threshold. gSAP retrieves up to four times as many positive pairs as the strongest AP surrogate at the same precision, and it degrades the least when queries with no positives in the database are added. Beyond thresholding, models trained with gSAP also learn better representations, with higher transfer, $k$NN and zero-shot classification accuracy.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09863 [cs.CV] |
| (or arXiv:2610.09863v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09863 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Giorgos Kordopatis-Zilos [view email]
[v1]
Wed, 7 Oct 2026 11:22:05 UTC (706 KB)
来源:arXiv:cs.LG · arxiv.org