arXiv:cs.CL· Jin Huang, Yutong Xie, Wanli Song, Xingjian Zhang, Walter Yuan, Matthew O. Jackson, Qiaozhu Mei·· 3 小时前AI 评分43
BehaviorBench:面向行为科学任务的基础模型基准测试
BehaviorBench: Benchmarking Foundation Models for Behavioral Science Tasks
AI 导读
研究者推出 BehaviorBench,从行为预测与模拟、策略决策、受试者特质推断、行为知识应用四项能力评估基础模型,并在个体与分布两个层面衡量表现。结果显示,该基准对通用 LLM 和用行为数据专门训练的行为基础模型仍具挑战性,个体层面与分布层面表现并不总是一致。通用 LLM 倾向低估人类回答的多样性,而行为基础模型在个体预测上常落后,基于多样行为数据微调可同时改善个体预测与分布对齐。
正文
Abstract:Foundation models have been increasingly applied to behavioral science domains such as psychology, sociology, and economics. While these models show promise in tasks such as survey response prediction and human-subject experiment simulation, there remains no systematic understanding of how well they perform across diverse behavioral science tasks. We introduce BehaviorBench, a comprehensive benchmark that evaluates foundation models along four core capabilities: (1) behavior prediction and simulation, (2) strategic decision-making, (3) subject-trait inference, and (4) behavioral knowledge application. Crucially, BehaviorBench evaluates model outputs at both the individual and distributional levels, capturing not only per-subject accuracy but also population-level alignment, an essential requirement for behavioral validity. Our evaluation shows that BehaviorBench remains challenging for leading general-purpose LLMs and behavior foundation models that are specifically trained with behavioral data. We find that individual-level and distributional performance do not always align. General-purpose LLMs tend to underestimate the diversity of human responses, whereas behavior foundation models often lag behind at individual-level prediction. Our investigation further demonstrates how fine-tuning on diverse behavioral data can improve both individual-level prediction and distributional alignment, balancing these two objectives. Our results highlight the importance of evaluation at both individual and distributional levels, establishing BehaviorBench as a foundation for developing and assessing behaviorally aligned AI systems. Our BehaviorBench and models can be accessed via this https URL
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2606.24162 [cs.CL] |
| (or arXiv:2606.24162v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24162 arXiv-issued DOI via DataCite |
Submission history
From: Jin Huang [view email]
[v1]
Tue, 23 Jun 2026 05:30:54 UTC (1,726 KB)
[v2]
Wed, 7 Oct 2026 17:05:41 UTC (1,108 KB)
来源:arXiv:cs.CL · arxiv.org