arXiv:cs.LG· Harry Lyu, Neil Thompson·· 6 小时前AI 评分53
O*NET-BENCH 审计 LLM judge:排序对了,职业级规模估错了
Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
AI 导读
Harry Lyu 与 Neil Thompson 发布 O*NET-BENCH,基于 45,796 条工人评分构建审计套件,在 4,501 条测试评分上评估 6 个模型家族的 33 种 judge 配置。
正文
Abstract:LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are acceptable, compared with 61.1% for occupation-matched workers. In one fine-tuned lineage, changing from pointwise scoring to a bundled few-shot/listwise protocol improves response ordering while reducing agreement with worker means at the task and occupation levels; this reversal replicates on a task- and worker-disjoint validation split under prespecified criteria. Cross-validated calibration largely removes mean bias, but calibrated scores explain at most 8.5% of individual worker-rating variance. Prediction-assisted estimation yields at most small precision gains at the studied label budgets. These results show that ranking agreement alone is insufficient for occupational measurement. Judges should be validated against the acceptance rates and aggregates their scores will be used to estimate.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Computers and Society (cs.CY); Machine Learning (cs.LG); General Economics (econ.GN) |
| Cite as: | arXiv:2610.02492 [cs.AI] |
| (or arXiv:2610.02492v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02492 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Harry Lyu [view email]
[v1]
Thu, 1 Oct 2026 21:14:49 UTC (330 KB)
来源:arXiv:cs.LG · arxiv.org