arXiv:cs.CL· Baowen Zhang, Wei Fan, Ruman Wang, Hangting Ye·· 6 小时前AI 评分41
LLM 在不完美表格问答中的错误理解研究
Understanding Errors in LLM-Based Question Answering over Imperfect Tables
AI 导读
基于 RADAR-T 的人类审核样本,研究对三种 LLM 的受控实验发现,重排行顺序会改变错误发现结果,即使表格内容与标准答案不变;LLM 在直接检查时,当含相关错误的行出现在表格靠后位置或彼此更集中时,更容易发现全部错误行。
正文
Abstract:Answering questions over imperfect tables requires handling errors that can affect the answer. We investigate two challenges for large language models (LLMs): whether error discovery depends on where errors appear in a table, and whether providing their locations is sufficient for accurate question answering (QA). Using human-reviewed instances from RADAR-T, we conduct controlled studies across three LLMs by varying row order and comparing original, error-marked, and repaired tables. First, reordering rows changes error discovery even when the table contents and gold answer remain unchanged. During direct inspection, LLMs are more likely to discover all rows containing relevant errors when these rows appear later in the table or are grouped more closely together. Second, providing verified error locations alone is insufficient for accurate QA: with code execution, accuracy on repaired tables exceeds that on error-marked tables by 39.0-59.1 percentage points across the three LLMs. As a practical application of these findings, we combine error discovery across shuffled table views with explicit guidance for verifying and handling the reported errors in a simple workflow, Geometry-Balanced Discovery and Intervention (GBDI). On RADAR-T, GBDI improves QA accuracy by 3.8-18.5 percentage points over a code-agent baseline across five LLMs (paired 95% confidence intervals exclude zero for four), at the cost of additional inference. These results highlight the importance of both reliable error discovery and effective error handling in QA over imperfect tables. Code is available at this https URL.
| Comments: | 41 pages, 7 figures |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.04687 [cs.CL] |
| (or arXiv:2610.04687v2 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.04687 arXiv-issued DOI via DataCite |
Submission history
From: Baowen Zhang [view email]
[v1]
Sat, 3 Oct 2026 18:03:41 UTC (496 KB)
[v2]
Tue, 6 Oct 2026 03:12:30 UTC (495 KB)
来源:arXiv:cs.CL · arxiv.org