跳到正文
arXiv:cs.CL· Hyojung Han, Jongmin Kim, Seung-Hun Jeon·· 3 小时前AI 评分36

无需标签的混合精度量化:用表示漂移为文本嵌入模型做训练后量化

Quantize by Drift: Label-Free Mixed-Precision Post-Training Quantization for Text Embedders

AI 导读

研究者提出一种无需相关性标签的混合精度训练后量化方法,用量化引起的表示漂移作为模块敏感度信号,在五个文本嵌入模型上以 macro Spearman 0.911 对采样混合精度方案排序。该方法在冻结方法和基线的三个嵌入模型上通过了预注册方向性假设(主预算下 3/3),但主预算下数值比同预算均匀精度低 0.99、0.85、1.01 分。输出漂移是稳健的粗粒度敏感度信号,而非普遍最优的分配目标。

正文

View PDF HTML (experimental)

Abstract:Mixed-precision post-training quantization needs a per-module sensitivity signal; for a text embedder the obvious one -- the retrieval quality a module costs when quantized -- needs relevance labels that deployments rarely have. We measure a label-free substitute: quantization-induced representation drift, obtained by quantizing one module, re-encoding the corpus, and recording how far the output embeddings moved from their full-precision positions. What is specific is the observable: the deployed output representation a dense retriever ranks with. Across five development embedders, configuration-level drift orders sampled mixed-precision plans against held-out retrieval quality at a macro Spearman of 0.911, the sensitivity transports across calibration corpora and retrieval domains in the usable regime, module drifts compose rank-consistently but not numerically, and relevance-derived sensitivity adds no consistent value. The method is one additive allocation under a hard packed-byte budget, with no labels and no search. On three embedders held untouched until method, baselines and hypotheses were frozen and sealed, the pre-registered directional hypothesis against the prior LieQ criterion holds (3/3 at the main budget, no collapse) and drift scores above a two-sided LieQ steelman in 2/3; but at the main budget drift is numerically lower than same-budget uniform precision on all three (-0.99, -0.85, -1.01 points), having reduced module and whole-model drift as designed. Output drift is thus a robust coarse sensitivity signal, not a universally optimal allocation objective: it avoids the catastrophic failures of the transferred signed-geometry adaptation and can remain usable at stressed budgets where uniform collapses, but fine-grained redistribution around a strong uniform operating point remains unresolved.
Comments: 26 pages, 22 tables, 4 figures
Subjects: Information Retrieval (cs.IR); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as: arXiv:2610.09227 [cs.IR]
  (or arXiv:2610.09227v1 [cs.IR] for this version)
  https://doi.org/10.48550/arXiv.2610.09227

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jong Min Kim [view email]
[v1] Tue, 6 Oct 2026 23:44:32 UTC (202 KB)

来源:arXiv:cs.CL · arxiv.org