跳到正文
arXiv:cs.LG· Xilin Wang, David Bau, Byron C. Wallace·· 4 小时前AI 评分39

如何训练你的模型生物:面向可解释性研究的模型生物多目标训练方法

How to train your model organism

AI 导读

针对可解释性研究中用于植入后门、谄媚等行为的模型生物,研究者提出应从目标行为植入、通用能力保持和输出自然性三个目标进行验证,并发现现有公开模型生物套件在训练后聊天质量与 CoT 自然性明显退化。为此他们提出基于模型融合的多目标训练方法,在临床推理人口统计偏差模型生物套件上,DPO 训练比监督微调更接近基座模型,该多目标方法也能更好保留能力与自然性。验证指标还能预测可解释性方法恢复植入行为的效果。

正文

View PDF HTML (experimental)

Abstract:Model organisms of alignment-relevant behaviors (e.g., backdoors, sycophancy, spurious correlations) have emerged as a key tool for evaluating whitebox interpretability techniques. We argue that the prevailing practice of training model organisms to a single objective of installing the target behavior is insufficient and propose validating model organisms with respect to three objectives with associated metrics: target-behavior installation, general-capability preservation (i.e., parametric knowledge, chat quality), and output naturalness (i.e., CoT and activations). We re-visit two publicly released organism suites using this validation framework and show that (1) chat quality and CoT naturalness degrade substantially across training recipes, and (2) validation metrics predict how well interpretability methods recover the installed behavior, e.g., a logit lens readout covaries with an organism's general capabilities. We introduce a multi-objective training approach based on model merging to train more realistic model organisms. Finally, on a new suite of model organisms targeting demographic biases in clinical reasoning, we compare training recipes and find that DPO training stays closer to the base model than supervised finetuning, and the proposed model optimization approach better preserves capabilities and naturalness. Auditing this suite with an investigator agent, we again observe validation metrics tracking bias recovery. In sum, training methods shape the interpretability conclusions an organism supports, and we argue that one should consider multiple objectives to draw generalizable conclusions about interpretability methods using (realistic) model organisms.
Subjects: Machine Learning (cs.LG)
Cite as: arXiv:2610.10203 [cs.LG]
  (or arXiv:2610.10203v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.10203

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Xilin Wang [view email]
[v1] Wed, 7 Oct 2026 15:02:57 UTC (1,185 KB)

来源:arXiv:cs.LG · arxiv.org