arXiv:cs.LG(机器学习,全量分类)· Sahil Kadadekar·· 14 小时前AI 评分44
安全原型并非安全方向:响应安全嵌入中的参考依赖与提示词混淆
A Safe Prototype Is Not a Safety Direction: Reference Dependence and Prompt Confounds in Response-Safety Embeddings
AI 导读
研究审计了近期一个 sleeper-agent 检测器提出的"与已知安全响应均值嵌入的余弦相似度"评分规则,发现原始正质心规则无法确定安全与不安全响应的分离方向。
正文
Abstract:Can response safety be scored by cosine similarity to the mean embedding of known-safe responses? A recent sleeper-agent detector proposes exactly this score, yet the raw positive-centroid rule is not identified: positive observations locate the safe class relative to an encoder origin, but do not determine which direction separates safe from unsafe responses. We audit the rule on two prompt-controlled, human-labeled corpora and one auxiliary jury-labeled source control, using four frozen encoders and prompt-grouped splits. On the human-labeled corpora the safe prototype reaches ROC-AUC 0.457-0.545, with two cells significantly below chance and one above, while an explicit safe-minus-unsafe reference reaches 0.588-0.738 on the same embeddings; on the jury control the prototype is inverted (0.358-0.405) and the reference reaches 0.754-0.793. At validation-calibrated 5% false-safe thresholds, the reference accepts more safe responses on PKU-SafeRLHF (0.153-0.263 versus 0.039-0.061 across encoders) and Aegis (0.189-0.291 versus 0.004-0.045), but not reliably on BeaverTails. A fully unlabeled held-out reference recovers part to most of the referenced ranking, much less when only 5% of the pool is unsafe, whereas 80-634 labeled unsafe responses recover most of it. Prompt-only ablations show that prompt-label composition can inflate uncontrolled evaluations. This is a bounded result about a raw positive centroid, not all one-class methods or safety-specialized guards. A class mean is a location, not necessarily a safety direction; a declared reference with enough unsafe mass identifies orientation.
| Comments: | Accepted at the NeurIPS 2026 Workshop on Foundations of Language Model Security (FLMSec). 15 pages, 3 figures, 11 tables. Code, results, and a verifier are in the ancillary files |
| Subjects: | Machine Learning (cs.LG); Computation and Language (cs.CL); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.01801 [cs.LG] |
| (or arXiv:2610.01801v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.01801 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sahil Kadadekar [view email]
[v1]
Thu, 1 Oct 2026 14:44:01 UTC (312 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org