arXiv:cs.CL· Samrajnee Ghosh, Ashish Goswami, Naman Agarwal, Hemanshu Garg, Chinmay Mittal, Mausam, Parag Singla·· 4 小时前AI 评分53
Percept-V 基准:多模态大模型在简单感知任务上表现远逊人类
The Percept-V Challenge: Can Multimodal LLMs Crack Simple Perception Problems?
AI 导读
研究者发布 Percept-V 数据集,包含 6000 张程序生成、无污染的图像,分为 30 个领域,对应 TVPS-4 框架中的视觉辨别等七种感知技能,推理和知识要求极低。实验显示 SoTA 专有和开源多模态大模型表现远低于人类,且图像中物体数量增多时性能快速下降;对开源模型微调有明显提升,但泛化到相关数据集的能力有限。论文已被 COLM 2026 接收。
正文
Abstract:Cognitive science research treats visual perception, the ability to understand and make sense of a visual input, as one of the early developmental signs of intelligence. Its TVPS-4 framework categorizes and tests human perception into seven skills such as visual discrimination, and form constancy. Do Multimodal Large Language Models (MLLMs) match up to humans in basic perception? Even though many benchmarks evaluate MLLMs on advanced reasoning and knowledge skills, there is limited research that focuses evaluation on simple perception. In response, we introduce Percept-V, a dataset containing 6000 program-generated uncontaminated images divided into 30 domains, where each domain tests one or more TVPS-4 skills. Our focus is on perception, so we make our domains quite simple and the reasoning and knowledge required for solving them are minimal. Since modern-day MLLMs can solve much more complex tasks, our a-priori expectation is that they will solve these domains very easily. Contrary to our belief, our experiments show a weak performance of SoTA proprietary and open-source MLLMs compared to very high human performance on Percept-V. We find that as the number of objects in the image increases, performance goes down rather fast. Our experiments also identify the perception skills that are considerably harder for all models. Fine-tuning an open-source MLLM shows considerable gains in performance, though the gains only marginally carry over to other related datasets, pointing to limitation in generalization abilities of the learned representations.
| Comments: | Accepted at COLM 2026 |
| Subjects: | Computation and Language (cs.CL); Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2508.21143 [cs.CL] |
| (or arXiv:2508.21143v4 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2508.21143 arXiv-issued DOI via DataCite |
Submission history
From: Samrajnee Ghosh [view email]
[v1]
Thu, 28 Aug 2025 18:22:38 UTC (507 KB)
[v2]
Wed, 8 Oct 2025 07:49:55 UTC (798 KB)
[v3]
Thu, 22 Jan 2026 08:36:54 UTC (2,205 KB)
[v4]
Fri, 2 Oct 2026 14:47:53 UTC (6,914 KB)
来源:arXiv:cs.CL · arxiv.org