跳到正文
arXiv:cs.LG· Euodia Dodd, Nata\v{s}a Kr\v{c}o, Igor Shilov, Matthew Wicker, Yves-Alexandre de Montjoye·· 4 小时前AI 评分42

无需参考模型估计模型级成员推理脆弱性

Estimating Model-Level Membership Inference Vulnerability Without Reference Models

AI 导读

研究者提出一种无需训练任何参考模型、仅凭目标模型训练与测试损失分布即可估计 LiRA 攻击模型级脆弱性的方法。

正文

View PDF HTML (experimental)

Abstract:Membership inference attacks (MIAs) have emerged as the standard tool for evaluating the privacy risks of AI models. However, state-of-the-art attacks require training numerous, often computationally expensive, reference models, limiting their practicality. We present a novel approach for estimating model-level vulnerability to the Likelihood Ratio Attack (LiRA), the strongest available attack, directly from the train and test loss distributions of the target model and without training any reference models. We show that LiRA's per-sample signal decomposes into a variance-ratio term and a residual mean-shift term, with the relative contribution of each determined by how much training collapses model uncertainty at the trained sample. This places models on a continuum, with different regimes calling for different reference-free loss-based statistics as proxies for LiRA TPR. The shapes of the loss distributions themselves indicate which proxy applies. We instantiate the framework with two natural proxies. At the heavy-tailed end, the LOSS attack TNR predicts LiRA TPR@FPR=$10^{-3}$ with RMSE 0.036 across 10 image classification architectures and 4 datasets, outperforming low-cost reference-model attacks such as RMIA. At the symmetric end, the LOSS attack AUC predicts LiRA TPR with RMSE 0.018 across five GPT-2 sizes from 10M to 1B parameters.
Subjects: Machine Learning (cs.LG); Cryptography and Security (cs.CR)
Cite as: arXiv:2510.19773 [cs.LG]
  (or arXiv:2510.19773v2 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2510.19773

arXiv-issued DOI via DataCite

Submission history

From: Euodia Dodd [view email]
[v1] Wed, 22 Oct 2025 17:03:55 UTC (425 KB)
[v2] Wed, 7 Oct 2026 15:12:23 UTC (1,062 KB)

来源:arXiv:cs.LG · arxiv.org