跳到正文
arXiv:cs.LG· Edward Chen, Yuntao Du·· 4 小时前

LLM 蒸馏推断初步研究:检测模型是否蒸馏自专有 LLM

Poster: A Preliminary Study of LLM Distillation Inference

AI 导读

研究者提出"蒸馏推断"方法,通过假设检验判断某模型是蒸馏自专有 LLM 还是独立训练,并用影子模型将嫌疑模型得分转换为校准 p 值。以 Qwen2.5-7B 为教师、Llama-3.2-3B 为嫌疑模型的初步实验中,该检验在 0.02 显著性水平下真阳性率达 1.0。该论文已被 CCS'26 接收为 poster 论文。

正文

View PDF HTML (experimental)

Abstract:Unauthorized model distillation, in which a model is trained on the outputs of a proprietary large language model (LLM), is a growing threat to model providers. We study distillation inference: determining whether a suspect model was distilled from another model or trained independently. We formulate this problem as a hypothesis test and estimate the behavior expected under each hypothesis by training shadow models: distilled shadow models learn from the teacher's reasoning traces, whereas independent shadow models learn only from reference answers. The auditor measures how closely each model predicts the teacher's reasoning outputs and then uses the shadow models to convert the suspect's score into a calibrated p-value. In a preliminary study using Qwen2.5-7B as the teacher and Llama-3.2-3B for the suspects, our test achieves a true positive rate of 1.0 at a significance level of 0.02. These results demonstrate the feasibility of using distillation inference to detect distillation attacks.
Comments: Accepted as a poster paper at the 2026 ACM SIGSAC Conference on Computer and Communications Security (CCS'26)
Subjects: Cryptography and Security (cs.CR); Machine Learning (cs.LG)
Cite as: arXiv:2610.12137 [cs.CR]
  (or arXiv:2610.12137v1 [cs.CR] for this version)
  https://doi.org/10.48550/arXiv.2610.12137

arXiv-issued DOI via DataCite (pending registration)

Related DOI: https://doi.org/10.1145/3830454.3846416

DOI(s) linking to related resources

Submission history

From: Yuntao Du [view email]
[v1] Thu, 8 Oct 2026 15:24:29 UTC (149 KB)

来源:arXiv:cs.LG · arxiv.org