arXiv:cs.LG· Krishna Kabra, Constantin Venhoff, Christian Schroeder de Witt·· 4 小时前AI 评分48
Weight Oracles:用语言模型直接读取神经网络权重
Weight Oracles: Reading Neural Network Weights with Language Models
AI 导读
研究者提出 Weight Oracles,即通过微调语言模型直接读取目标网络原始权重来诊断其属性,无需行为测试。该方法在小型 Transformer 上模拟前向传播达到 99% 的留出准确率,并在未见过的后门检测中取得 AUROC 0.93(注意力路由后门)和 0.81(多样化威胁分布)。其扩展到现实模型规模仍是主要未解挑战。
正文
Abstract:Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: through a staged curriculum and an external chain-of-computation that delegates parameter-free operations to deterministic code, an explainer LLM learns to simulate the forward pass of small transformers from their weights, achieving 99% holdout accuracy on unseen targets. Phase II repurposes this infrastructure for safety auditing. We train an oracle on natural language diagnostic questions about weight anomalies using only benign pathologies as training signal, and evaluate it zero-shot on backdoors absent from training. The oracle achieves AUROC 0.93 on attention-routed backdoors and 0.81 across a diversified threat distribution including stealth and adversarially regularized variants. Hand-crafted statistical detectors are sharp on the threat models they implicitly target but collapse on threat-model shift, while the oracle remains uniformly competent across attack types. Scaling to realistic model sizes remains the principal open challenge.
| Comments: | Spotlight at the NeurIPS 2026 Workshop on Neural Network Artifacts as a New Data Modality (NeuralArtifacts), Paris. 14 pages, 10 figures |
| Subjects: | Machine Learning (cs.LG); Cryptography and Security (cs.CR) |
| Cite as: | arXiv:2610.07334 [cs.LG] |
| (or arXiv:2610.07334v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.07334 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Krishna Kabra [view email]
[v1]
Mon, 5 Oct 2026 20:08:29 UTC (833 KB)
来源:arXiv:cs.LG · arxiv.org