AVMeme Exam:面向 LLM 语境与文化知识及思考能力的多模态多语言多文化基准
AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
研究者推出 AVMeme Exam,一个由人工筛选的超过一千个互联网标志性声音与视频组成的基准,覆盖语音、歌曲、音乐与音效,每个 meme 配有从表层内容到语境、情感、用法与世界知识的问答及元数据。系统评测显示,当前多模态大语言模型在无文本音乐和音效上表现不佳,在语境与文化层面的思考也弱于表层内容理解。该基准已被 COLM 2026 接收。
Authors:Xilin Jiang, Qiaolin Wang, Junkai Wu, Xiaomin He, Zhongweiyang Xu, Yinghao Ma, Minshuo Piao, Kaiyi Yang, Xiuwen Zheng, Riki Shimizu, Yicong Chen, Arsalan Firoozi, Gavin Mischler, Sukru Samet Dindar, Richard Antonello, Linyang He, Tsun-An Hsieh, Xulin Fan, Yulun Wu, Yuesheng Ma, Chaitanya Amballa, Weixiong Chen, Jiarui Hai, Ruisi Li, Vishal Choudhari, Cong Han, Yinghao Aaron Li, Adeen Flinker, Mounya Elhilali, Emmanouil Benetos, Mark Hasegawa-Johnson, Romit Roy Choudhury, Nima Mesgarani
Abstract:Internet audio-visual clips convey meaning through time-varying sound and motion, which extend beyond what text alone can represent. To examine whether AI models can understand such signals in human cultural contexts, we introduce AVMeme Exam, a human-curated benchmark of over one thousand iconic Internet sounds and videos spanning speech, songs, music, and sound effects. Each meme is paired with a unique Q&A assessing levels of understanding from surface content to context and emotion to usage and world knowledge, along with metadata such as original year, transcript, summary, and sensitivity. We systematically evaluate state-of-the-art multimodal large language models (MLLMs) alongside human participants using this benchmark. Our results reveal a consistent limitation: current models perform poorly on textless music and sound effects, and struggle to think in context and in culture compared to surface content. These findings highlight a key gap in human-aligned multimodal intelligence and call for models that can perceive contextually and culturally beyond the surface of what they hear and see. Project page: this http URL
| Comments: | Accepted by COLM 2026; this http URL |
| Subjects: | Sound (cs.SD); Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV); Multimedia (cs.MM); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2601.17645 [cs.SD] |
| (or arXiv:2601.17645v2 [cs.SD] for this version) | |
| https://doi.org/10.48550/arXiv.2601.17645 arXiv-issued DOI via DataCite |
Submission history
From: Xilin Jiang [view email]
[v1]
Sun, 25 Jan 2026 01:40:15 UTC (29,905 KB)
[v2]
Mon, 5 Oct 2026 21:00:58 UTC (29,802 KB)
来源:arXiv:cs.CL · arxiv.org