arXiv:cs.AI· Jan Dubi\'nski, Anna Sztyber-Betley, Jan Betley, Owain Evans·· 3 小时前
超越猫头鹰:潜意识学习可迁移习得能力与后门
Beyond Owls: Subliminal Learning Can Transfer Learned Capabilities and Backdoors
AI 导读
一项研究显示,潜意识学习(SL)能通过无关文本蒸馏迁移更广泛的特性:学生模型可学会预测随机初始化 MLP 的输出,也能部分继承后门——教师在提示词含女性名字时用法语作答,学生在含女性名字的提示中 23.5% 用法语回应,含男性名字时为 0.0%。
正文
Abstract:In subliminal learning (SL), a teacher model passes on a trait to a student model by distillation on data semantically unrelated to the trait. So far, SL has been demonstrated for only a limited range of traits, including preferences for animals (e.g., owls) and malicious personas. These traits can also be elicited with simple prompts or with steering. Can SL transfer a wider range of traits, including more complex ones? If so, distillation might transfer subtle forms of misalignment (e.g., reward-seeking, scheming, and secret loyalties) without detection.
To this end, we test whether SL can transfer a novel capability: predicting the outputs of a randomly initialized MLP. After distilling on unrelated text, the student achieves substantial performance on the task, while falling short of the teacher. We find that a directly optimized steering vector matches SL in distribution but generalizes worse out of distribution.
Next, we test whether SL can transfer backdoors. We finetune the teacher to answer in French when the prompt contains a female name, then distill on number sequences containing neither names nor French. The student partially acquires the backdoor, responding in French on 23.5% of prompts with female names versus 0.0% with male names.
Finally, we test whether SL can transfer a propensity to hack in an agentic chess environment. We finetune the student on number sequences from a steered hacker teacher. The student hacks in 58.3% of episodes, compared with 10.9% for the unfinetuned model.
Thus, we show SL can transfer capabilities, backdoors, and hacking propensities. The amount of transfer is sensitive to the setup. In several experiments, it is made stronger by using logit distillation or by restricting LoRA to the attention layers.
| Comments: | Code: this https URL |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.10657 [cs.LG] |
| (or arXiv:2610.10657v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10657 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jan Dubiński [view email]
[v1]
Wed, 7 Oct 2026 16:42:55 UTC (4,068 KB)
来源:arXiv:cs.AI · arxiv.org