arXiv:cs.LG· Erik Skalnes, Layne C. Price, Raviteja Anantha, Michael Oberst·· 4 小时前
用干预迁移评估评分标准生成
Evaluating Rubric Generation with Interventional Transfer
AI 导读
研究者提出 Interventional Transfer(IT)方法,用于规模化评估 LLM 生成的评分标准质量,其核心思路是:若一份回答被扰动为通过/不通过某条标准时,两份评分标准同步变化,则二者相似。
正文
Abstract:Instance-specific rubrics are common in AI benchmarks where reliable evaluation requires specific expert knowledge. This approach is difficult to scale, prompting research into the generation of rubrics with large language models (LLMs). However, even when expert rubrics are available as references, it is unclear how to productively evaluate the quality of generated rubrics at scale. In this paper, we introduce a method for the evaluation of rubric generation, which we call Interventional Transfer (IT), based on the idea that two rubrics are similar if they move together when a response is perturbed to pass/fail one of them. In contrast to existing approaches for evaluation of rubric generation, we argue that different forms of interventional transfer can be used to evaluate the utility of generated rubrics for different tasks. For instance, we apply this approach in a case study on HealthBench, where we demonstrate an asymmetry in rubrics generated by Qwen3.8-27B, Deepseek-V4-Flash, and Opus-5, used to evaluate responses from GPT-5.6-Terra. Perturbations that degrade responses according to the generated/expert rubric transfer into lower scores on the corresponding expert/generated rubric, but perturbations that improve on one rubric do not reliably transfer into higher scores on the other. We argue that this finding has implications for the usage of LLM-generated rubrics for performance monitoring and hill-climbing. We contrast our approach with existing approaches for rubric evaluation, which do not surface the same asymmetry that we observe.
| Comments: | 29 pages, 6 figures. Code and cached model outputs: this https URL |
| Subjects: | Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.10809 [cs.LG] |
| (or arXiv:2610.10809v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10809 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Erik Skalnes [view email]
[v1]
Wed, 7 Oct 2026 19:12:45 UTC (42 KB)
来源:arXiv:cs.LG · arxiv.org