跳到正文
arXiv:cs.AI· Nina Nusbaumer, Iria de-Dios-Flores, Corentin Bel, Christophe Pallier, Guillaume Wisniewski, Beno\^it Crabb\'e·· 7 小时前AI 评分31

STRUCTURALCOST:用于建模人类句子处理难度的控制性阅读时间数据集

STRUCTURALCOST: A controlled reading time dataset for modeling human sentence processing difficulty

AI 导读

研究者发布 STRUCTURALCOST 自定步速阅读数据集,覆盖 475 名参与者、40,800 条观测,用于分离长距离主谓依存解析的处理成本,该工作将在 EMNLP 2026 发表。结果显示主句动词处阅读时间随依存距离增加,且由句法嵌套而非线性距离驱动;n-gram、SSM 和 Transformer 等语言模型仅部分反映这一难度曲线,普遍低估人类的整合成本,且该差距跨架构与模型规模持续存在。

正文

View PDF HTML (experimental)

Abstract:We introduce STRUCTURALCOST, a self-paced reading dataset of 475 participants and 40,800 observations isolating the processing cost of long-distance subject-verb dependency resolution. We replicate a low-powered psycholinguistic finding at NLP scale, namely that human reading times at the main verb increase with dependency length, driven by syntactic embedding beyond linear distance. Different language models -- spanning n-gram models, SSMs, and transformers -- partially mirror this graded difficulty profile, yet underestimate the integration cost humans incur, with a gap that persists across architectures and model sizes. This suggests these models capture the predictive component of human processing but not the full integration cost that working memory imposes. STRUCTURALCOST provides data needed to drive progress toward evaluating the cognitive plausibility of language models.
Comments: Will be published at EMNLP 2026
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.08208 [cs.CL]
  (or arXiv:2610.08208v1 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2610.08208

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Nina Nusbaumer [view email]
[v1] Tue, 6 Oct 2026 12:00:36 UTC (100 KB)

来源:arXiv:cs.AI · arxiv.org