arXiv:cs.CL· Jiayu Feng·· 3 小时前AI 评分40
LLM 回答州政策出错时是在跨州替换吗?一项针对美国 51 个司法辖区的测量研究
Right Number, Wrong State? Measuring Cross-Jurisdiction Substitution in LLM Recall of State Policy
AI 导读
研究用固定措辞、只更换辖区的极简设计,测试 Claude Sonnet 5.5 与 GPT-5.6 Sol 对美国 51 个司法辖区三项 Medicaid 收入资格数值的召回,两者在 153 个项目中分别有 10 和 25 项可复现地给出其他州的当前数值,且两次独立重复完全一致。
正文
Abstract:When an LLM answers a state-specific policy question wrongly, it may be hallucinating, or it may be returning a real value that holds in another state. We test this with a minimal-set design: the question wording is fixed and only the jurisdiction varies, across the 50 U.S. states and the District of Columbia (51 jurisdictions) and three exactly defined Medicaid income-eligibility quantities. Gold values come from an official data book and agree with an independent source in 101 of 102 checked cells. Under a pre-registered protocol, Claude Sonnet 5.5 and GPT-5.6 Sol reproducibly give another state's current value, identical across two independent repeats, for 10 and 25 of 153 items. Attribution is fragile, however. Crediting any wrong answer that equals another state's value yields 3-5x more reproducible substitutions than checking every number in the asked state's own records, because many apparent cross-state answers are the asked state's own values under another convention or from an earlier year. Claims about cross-jurisdiction error need a complete same-state reference set. We will release the protocol, gold table, and all model outputs.
| Comments: | 6 pages, 3 figures, 1 table |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.09458 [cs.CL] |
| (or arXiv:2610.09458v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09458 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jiayu Feng [view email]
[v1]
Wed, 7 Oct 2026 05:12:42 UTC (146 KB)
来源:arXiv:cs.CL · arxiv.org