arXiv:cs.LG· James Mickens·· 4 小时前AI 评分41
语言不可读性对 LLM 安全的影响
The Implications of Linguistic Illegibility for LLM Security
AI 导读
研究提出"语言不可读性"概念,指 LLM 的外部语言输出与机制提取的语言特征无法可靠反映模型内部计算,因为其内部运算是激活空间上的数学而非语言。因此依赖模型语言自报告的链式思维监控、宪法式自我批评、激活探针等安全机制永远无法完全可靠。作者主张用污点追踪观察模型输出,并配合稳健虚拟化、第三方审计等沙箱机制,可缓解前沿模型的沙箱漏洞利用。
正文
Abstract:LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks. We argue that the specter of linguistic illegibility is unavoidable for LLMs whose internal computations are not directly expressed via language, but rather math over activation spaces (with lossy translations between activation spaces and natural language happening at the bookends). If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors) can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all. We argue that observing a model's outputs using taint tracking is a promising approach for an effective sandbox: regardless of how a model linguistically self-reports, a taint tracking policy can define, a priori, various pieces of system state that should never be influenced by model-produced data. We also discuss several additional sandboxing mechanisms (e.g., robust virtualization, third-party auditing of sandboxing configurations) which collectively provide a critical floor beneath linguistic monitoring, and would have mitigated recent sandbox exploits by frontier models.
| Subjects: | Machine Learning (cs.LG); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2609.02852 [cs.LG] |
| (or arXiv:2609.02852v2 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.02852 arXiv-issued DOI via DataCite |
Submission history
From: James Mickens [view email]
[v1]
Wed, 2 Sep 2026 17:37:22 UTC (33 KB)
[v2]
Tue, 6 Oct 2026 21:19:03 UTC (34 KB)
来源:arXiv:cs.LG · arxiv.org