arXiv:cs.CL· Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)·· 6 小时前AI 评分27
Jev 作为金融证据判定器的压力测试:同数引用替换评估
Same-Number Citation Swaps: Stress-Testing Jev as a Financial Evidence Judge
AI 导读
研究以 Jev 作为 GPT-4.1-mini 计算轨迹的来源支撑验证器,用"同数引用替换"方法在固定操作数与算术的前提下移动引用位置,检验概率性证据验证在数字匹配之外的增量价值。
正文
Authors:Chuhong Xu (Sofia University), Bo Su (Indiana University), Ziyao Chen (University of California, San Diego), Ruiyang Xu (Northeastern University), Shimeng Dai (Michigan State University), Xinyu Qiu (Northeastern University)
Abstract:Financial reports repeat values across periods, metrics and accounting lines, allowing an LLM-generated calculation to be numerically correct while citing the wrong financial role. We evaluate what probabilistic evidence verification adds beyond number matching using Jev as a source-support verifier for GPT-4.1-mini calculation traces. A signed-number-at-pointer baseline explains most recovery over exact quotation checks. To isolate the remaining role-recognition problem, we hold operands and arithmetic fixed, move citations between same-number cells, and retain controls that express equivalent facts. These contrasts reveal both wrong-role citations that pass and valid alternative citations that are withheld. Explicit column labels improve selected wrong-role decisions while also lowering support for some equivalent evidence. A constructed follow-up on 36 new source pages, labeled by a non-author reviewer, extends this evaluation and exposes the same tradeoff between detecting role errors and retaining valid citations. The contribution is a controlled evaluation that identifies what a probabilistic financial verifier distinguishes when numerical matching is held fixed. For LLM-based financial assistants, it makes numerical correctness, cited-role support and acceptance outcomes separately assessable.
| Comments: | counterfactual citation perturbation, evidence attribution verification, financial document question answering, Jev, LLM-as-a-judge, probabilistic source verification, tabular numerical reasoning |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.08675 [cs.CL] |
| (or arXiv:2610.08675v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08675 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinyu Qiu [view email]
[v1]
Tue, 6 Oct 2026 16:57:47 UTC (90 KB)
来源:arXiv:cs.CL · arxiv.org