跳到正文
arXiv:cs.LG· Jeremy Canale·· 7 小时前AI 评分51

arXiv 论文提出对齐缩放定律框架并给出首批预注册测量

Toward Alignment Scaling Laws: A Framework and First Preregistered Measurements

AI 导读

arXiv 论文(arXiv:2610.08540)将对齐建模为一族可测缩放关系,负担 B_r(N)=a_r N^alpha_r,alpha_r<1 表示缩放有帮助,>1 则累积对齐债务。

正文

View PDF HTML (experimental)

Abstract:Whether alignment gets easier or harder as models grow is often argued from isolated findings, as if alignment were one property. We treat it as a family of measurable scaling relations: for each risk category r, the alignment burden needed to hold a fixed safety target is modeled as B_r(N)=a_rN^alpha_r, with N a capability proxy; against a budget proportional to N, scaling helps if alpha_r<1, keeps pace if alpha_r~1, and accumulates alignment debt if alpha_r>1. We give three operationalizations of burden and distinguish observed, audited and true alignment. A toy model, in which corrections consume capability headroom, makes the consequences explicit. We prove that the largest exponent among corrected risks, not an average, sets the long-run regime; that above 1 any policy holding headroom above a floor must grow super-exponentially; that, for burdens that are positive mixtures of power laws, fits on small models underestimate large-scale exponents; and that an audit that uncovers hidden failures without false positives never underestimates true alignment. We propose a pre-registrable protocol and apply reduced versions of it twice. A preregistered reanalysis of public adversarial-training data for Pythia classifiers finds that the compute needed to bring attack success under 10% grows as N^0.60. A preregistered pilot on Qwen2.5 0.5B-72B finds exponents of -0.05 for truthfulness and 0.48 for stated dispositions (both scaling helps under its reduced rule, though local slopes approach 1 at the top; replicated on Qwen3 0.6B-14B), while sycophancy (0.89, or 0.83 with two seeds added at 72B) and a planted backdoor are undetermined: the backdoor is removed quickly when its trigger is known but survives blind safety training at four of five sizes. We release four browser games that play these laws (this http URL). We make no claim about which regime holds for current frontier models.
Comments: 34 pages, 24 figures, 8 tables. Games: this https URL. Preregistrations: this https URL, this https URL, this https URL
Subjects: Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.08540 [cs.AI]
  (or arXiv:2610.08540v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.08540

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jeremy Canale [view email]
[v1] Tue, 6 Oct 2026 15:29:07 UTC (2,025 KB)

来源:arXiv:cs.LG · arxiv.org