跳到正文
arXiv:cs.CL· Zhaohui Geoffrey Wang·· 3 小时前AI 评分45

研究揭示 LLM 后训练中 RankMe 监控的失效:秩上升反而预示表示退化

When Rank Rises as LLMs Degrade

AI 导读

针对 Qwen3-0.6B 的四类退化模式与三个随机种子的对照研究表明,数据重复会使留出损失相对健康状态恶化 75%,同时原始与中心化 RankMe 均上升,后者变化达 13.5 个合并标准差,协方差有效秩升至健康值近两倍。这种谱色散而非坍缩的失效会让单侧监控把最差检查点判为最健康;仅靠重新校准无法确定方向,需用双侧多通道序列监控并配合留出健康数据。

正文

View PDF HTML (experimental)

Abstract:Post-training adapts language models in non-stationary environments. Practitioners monitor representation health with RankMe and related spectral statistics, often assuming that rank falls when representations degrade. We show that this assumption is unsafe for LLM post-training. In a controlled study of Qwen3-0.6B with four degradation modes and three seeds, data duplication worsens held-out loss by 75% relative to healthy while increasing both original and centred RankMe; the latter changes by 13.5 pooled standard deviations. Covariance effective rank rises to nearly twice its healthy value. This failure is spectral dispersion rather than collapse, so a one-sided monitor rates the worst checkpoint as the healthiest. By contrast, a learning-rate misconfiguration lowers centred RankMe and k95, while uncentred RankMe is inconsistent across seeds. Direction is therefore a property of the regime-statistic pair and cannot be fixed by recalibration alone. We also distinguish two often-conflated statistics: RankMe normalises singular values, whereas covariance effective rank normalises eigenvalues. On raw intermediate-layer states in the pretrained model, massive activations pin the latter near 1 out of dimension d while RankMe retains usable range. We then test a two-sided, multichannel sequential monitor with separate calibration and test data. In a pre-registered shared-prefix, leave-one-seed-out evaluation, it detects all three damage regimes in every fold 10 to 60 steps after the fork and separates dispersion from downward-rank damage by firing direction. However, it never precedes held-out probe loss, and calibration with two seeds produces false alarms on the held-out healthy seed. Spectral monitoring can diagnose failure regimes, but it does not warn earlier than held-out loss, and validity claims require held-out healthy data.
Comments: NeurIPS 2026 Workshop on Continual Learning for Foundation Models and Agents (CL4FMAgents); 8 pages + appendix
Subjects: Machine Learning (cs.LG); Computation and Language (cs.CL)
Cite as: arXiv:2610.09647 [cs.LG]
  (or arXiv:2610.09647v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2610.09647

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Zhaohui Wang [view email]
[v1] Wed, 7 Oct 2026 08:22:30 UTC (54 KB)

来源:arXiv:cs.CL · arxiv.org