arXiv:cs.CL· Tao Shi, Chaoyi Xiang, Qiongkai Xu, Jey Han Lau·· 4 小时前AI 评分44
OLMo-Detect:面向大语言模型成员推理的多阶段、混杂因素受控基准
OLMo-Detect: A Multi-Stage, Confounder-Controlled Benchmark for Membership Inference on Large Language Models
AI 导读
基于全开放 OLMo 2 流程构建的 OLMo-Detect 基准覆盖预训练、中训练与后训练三个阶段,并用 infini-gram 严格过滤非成员样本。在 OLMo 2 系列上评测 15 种无监督与 3 种有监督成员推理攻击,最佳 AUC 仅 0.68,且无监督方法在分布偏移下 AUC 最多变化 0.42。
正文
Abstract:Membership inference on large language models (LLMs) aims to determine whether a given text sample was included in an LLM's training data, without access to its training corpus. Despite recent progress, existing benchmarks suffer from three limitations: limited coverage of training stages, insufficient distributional alignment between members and non-members, and lack of rigorous filtering of non-members against the training corpus. To address these limitations, we propose OLMo-Detect, a multi-stage, confounder-controlled benchmark built upon the fully open OLMo 2 pipeline. OLMo-Detect spans pre-training, mid-training, and post-training, explicitly aligns members and non-members on three key axes, and rigorously filters non-members via infini-gram. To assess robustness to distribution shifts, we further introduce OLMo-Detect (Shifted), a variant where members are misaligned with non-members. We evaluate 15 unsupervised and 3 supervised membership inference attacks (MIAs) across the OLMo 2 family, finding that: (i) overall performance is limited: the best unsupervised and supervised MIAs both reach an AUC of only 0.68, and supervised MIAs degrade under cross-domain evaluation; (ii) MIA performance peaks at mid-training and is lower at pre-training and post-training, a pattern driven by data type rather than a stage effect: curated math data is far more detectable than other types; (iii) overall scores improve from 1B to 13B but plateau at 32B; and (iv) no unsupervised MIA is robust to distribution shifts, with AUCs shifting by up to 0.42. Finally, we find that our findings on OLMo 2 generalize to OLMo 3 and non-OLMo models.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.02986 [cs.CL] |
| (or arXiv:2610.02986v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02986 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tao Shi [view email]
[v1]
Fri, 2 Oct 2026 08:22:46 UTC (135 KB)
来源:arXiv:cs.CL · arxiv.org