跳到正文
arXiv:cs.LG· Hamidreza Saghir·· 5 小时前AI 评分38

语言模型轨迹中的 OOD 检测:有用特征为何给出反向评分

Useful Features, Backward Scores: OOD in Language-Model Trajectories

AI 导读

研究分析语言模型轨迹中 OOD 检测的特征区分能力与异常排序之间的落差:在 Spam 数据上,D²HScore 的输入适配版本经长度匹配后 AUROC 从 0.919 降至 0.530。

正文

View PDF HTML (experimental)

Abstract:Out-of-distribution (OOD) detectors prioritize inputs for closer inspection. Yet features that distinguish input groups need not yield a useful anomaly ranking. We analyze this gap in language-model trajectories under text-length control and fixed score directions. On Spam development data, an input adaptation of D^2HScore falls from raw AUROC 0.919 to 0.530 after length matching. On length-matched, held-out HateSpeech inputs, the same features yield AUROC 0.644 for a labeled linear classifier but 0.444 for an ID-fitted distance score. ToxicChat shows the same contrast. Feature-selection and backbone controls retain the main reversal pattern. Frozen Civil Comments and TweetEval irony tests also reverse (0.467 and 0.435), extending the finding beyond toxicity. In these contrasts, anomalous groups have farther centers but tighter spread. A labeled, fixed-center feature-space intervention changes rankings: equalizing spread helps some tasks and harms others. OOD evaluation must check the chosen score's ranking even when its features distinguish the classes.
Subjects: Computation and Language (cs.CL); Machine Learning (cs.LG)
ACM classes: I.2.7; I.2.6
Cite as: arXiv:2605.00269 [cs.CL]
  (or arXiv:2605.00269v2 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2605.00269

arXiv-issued DOI via DataCite

Submission history

From: Hamidreza Saghir [view email]
[v1] Thu, 30 Apr 2026 22:06:02 UTC (91 KB)
[v2] Fri, 2 Oct 2026 06:47:11 UTC (669 KB)

来源:arXiv:cs.LG · arxiv.org