跳到正文
arXiv:cs.CL· Jin Huang, Diego Ferreras Garrucho, Yutong Xie, Walter M. Yuan, Qiaozhu Mei, Chen Lian, Jonathon Hazell·· 6 小时前AI 评分43

HouseholdBench:评估大语言模型作为家庭经济行为预测器

HouseholdBench: Evaluating Large Language Models as Predictors of Household Economic Behavior

AI 导读

研究团队推出 HouseholdBench,整合 6 项美国住户调查与 32 个预测任务,覆盖消费、收入、劳动、预期和住房等数值、分类及概率型结果,评估 13 个专有与开源 LLM 对家庭行为的预测及政策响应能力。

正文

View PDF HTML (experimental)

Abstract:Large language models (LLMs) have the potential to meet a key goal in economics: a quantitative model of household decision making, across a variety of settings. Yet existing evaluations cover few surveys and outcomes, and do not study how households adjust to changing economic conditions. We introduce a new evaluation, HouseholdBench, which unites 6 U.S. household surveys and 32 prediction tasks spanning numeric, categorical and probabilistic outcomes, related to consumption, income, labor, expectations, and housing. Using past behavior, demographics and macroeconomic conditions, the tasks test whether LLMs predict behavior, including how households adjust to changes in various policies. We evaluate 13 proprietary and open-weight LLMs against a no-change baseline and a gradient-boosted tree model. Most LLMs outperform the no-change baseline, including for policy response tasks -- with the best model lowering error for numeric outcomes by 12.2%. Across most tasks, gradient-boosted trees rank first; leading proprietary LLMs approach their performance, but open-weight models lag. LLMs exhibit systematic over- and underprediction across different tasks. We identify methods that enable a 4 billion parameter open-weight model to match proprietary models' performance: fine-tuning and aggregating 16 predictions per observation. Improvements generalize to policy-response tasks, which are excluded from fine-tuning. We release our datasets, code, and leaderboard on our website: this https URL
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2610.07563 [cs.CL]
  (or arXiv:2610.07563v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.07563

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Jin Huang [view email]
[v1] Tue, 6 Oct 2026 00:50:28 UTC (549 KB)

来源:arXiv:cs.CL · arxiv.org