跳到正文
arXiv:cs.CL· Nataliya Stepanova, Ivan Titov, Emily Allaway, Bj\"orn Ross·· 7 小时前AI 评分35

LLM 刻板印象评估缺失的最小对比对

The Missing Minimal Pair: Stereotype Evaluation in LLMs

AI 导读

研究指出,仅比较一对对比性刻板印象句子的对数似然往往不可靠,改写同一刻板印象的属性即可产生逻辑不一致的偏好。作者提出双最小对比对评估框架,通过数据增强生成释义与替换属性,覆盖英语、俄语、西班牙语和中文刻板印象,并设计两项评估指标,其中基于互信息(MI)的指标更适合聚合,可跨语言、跨模型更稳健地比较刻板印象强度。代码已公开。

正文

View PDF HTML (experimental)

Abstract:A common approach to measuring bias in Large Language Models is to compare the log-likelihoods of two contrastive stereotype sentences. We argue that such single-pair comparisons are often unreliable: simply rewriting the same stereotype with an alternative attribute can yield logically inconsistent preferences. To address this, we propose a dual minimal pair setup that introduces two axes of comparison for robust stereotype evaluation. First, we present a data-augmentation framework that fills critical gaps in existing stereotype datasets by generating paraphrases and alternate attributes. We apply our framework on a set of English, Russian, Spanish and Chinese stereotypes. Second, we introduce two evaluation metrics tailored to the dual minimal pair setup. One of these metrics provides a new perspective on bias by modeling the mutual information (MI) between social groups and stereotyped attributes. This MI-based metric is better suited for aggregation and enables more robust comparisons of stereotype strength across different languages and models.
Our code is available at this https URL.
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.08747 [cs.CL]
  (or arXiv:2610.08747v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08747

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nataliya Stepanova [view email]
[v1] Tue, 6 Oct 2026 17:41:44 UTC (2,234 KB)

来源:arXiv:cs.CL · arxiv.org