arXiv:cs.AI· Phillip Howard, Xin Su, Kathleen C. Fraser·· 5 小时前AI 评分44
大型视觉语言模型中的跨文化价值观归因研究
Cross-Cultural Value Attribution in Large Vision-Language Models
AI 导读
研究分析了大型视觉语言模型(LVLM)在宗教、国籍和社会经济地位等文化语境下的价值观判断偏差。通过 480 万次 LVLM 生成,研究识别出三种可跨多种架构模型复现的调研锚定偏差模式,并发现国籍锚定以文本为主,宗教和社会经济锚定则高度依赖图像。该工作已被 EMNLP 2026 Findings 接收。
正文
Abstract:The rapid adoption of large vision-language models (LVLMs) in recent years has been accompanied by growing fairness concerns due to their propensity to reinforce harmful societal stereotypes. While significant attention has been paid to such fairness concerns in the context of social biases, relatively little prior work has examined the presence of stereotypes in LVLMs related to cultural contexts such as religion, nationality, and socioeconomic status. In this work, we aim to narrow this gap by investigating how LVLM judgments about a person's moral, ethical, and political values vary across cultural contexts presented in images. We conduct a multi-dimensional analysis of such value judgments in popular LVLMs using counterfactual image sets, which depict the same person across different cultural contexts. Our evaluation framework pairs descriptive analyses (Moral Foundations Theory categorization, lexical analyses, and value sensitivity) with a novel grounding analysis that compares LVLM cross-context variation against two large-scale human surveys (MFQ-2 and WVS Wave 7). Across 4.8 million LVLM generations, we identify three survey-grounding bias patterns that replicate across multiple architecturally diverse models. Additional ablations show that nationality grounding is text-dominant while religion and socioeconomic grounding depend strongly on the image, and that image conditioning can amplify survey-grounding bias patterns.
| Comments: | Accepted to EMNLP 2026 Findings |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2604.09945 [cs.CV] |
| (or arXiv:2604.09945v3 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2604.09945 arXiv-issued DOI via DataCite |
Submission history
From: Phillip Howard [view email]
[v1]
Fri, 10 Apr 2026 22:53:41 UTC (4,324 KB)
[v2]
Thu, 2 Jul 2026 01:29:14 UTC (21,589 KB)
[v3]
Mon, 5 Oct 2026 22:09:32 UTC (21,600 KB)
来源:arXiv:cs.AI · arxiv.org