跳到正文
arXiv:cs.AI· Saanvi Khetan, Sankar Balasubramanian·· 6 小时前AI 评分61

arXiv 论文:LLM 在财务建议缺失事实时以身份先验替代,Llama-3.1-8B-Instruct 上身份差距从 4.78 升至 10.34 个百分点

Thin Evidence, Thick Priors: How Language Models Substitute Identity for Missing Financial Facts

AI 导读

论文测试 Llama-3.1-8B-Instruct 在财务建议中如何处理缺失信息:基于 100 个财务画像、138 个人设和七种披露条件共 96600 条提示,两个财务完全相同的人设之间推荐股票配置差距从完全披露时的 4.78 个百分点升至无财务事实时的 10.34 个百分点,聚类自助法估计比值为 2.16(95% 区间 1.69 至 2.79)。

正文

View PDF HTML (experimental)

Abstract:People increasingly ask large language models what to do with their money, yet seldom describe their finances in full. This paper asks what a model does with the gap. Holding finances fixed and changing only who the investor is said to be, we grade the financial evidence in the prompt from eight facts to none and measure how far the recommended equity allocation moves. Across 96,600 prompts to Llama-3.1-8B-Instruct, built from 100 financial profiles, 138 personas and seven disclosure conditions, the average gap between two personas with identical finances rises from 4.78 percentage points at full disclosure to 10.34 points with no financial facts. A two-way cluster bootstrap counting duplicated prompts once places the ratio at 2.16 (95% interval 1.69 to 2.79), and the rise is already 1.69-fold with a single fact left. Identity explains 5% of within-profile variation in advice at full disclosure and 96% with no disclosure. Household size is the only attribute whose influence grows reliably as evidence is withdrawn. Once standard errors are clustered on the persona, the unit to which identity was assigned, most attribute-specific interactions reported in the conference version lose significance, and gender instead appears as a small standing gap that full disclosure does not close. Stating risk appetite alone brings the swing into the range seen with two to seven generic facts. With no facts, the model's one-line rationale cites incomes, debts and savings it was never told, and these invented finances turn adverse more often for larger households. Inside the network, gender is linearly decodable at every layer, and ablating the gender direction at five layers leaves the aggregate identity swing unchanged. Advisory systems built on such models should be audited at the disclosure levels users actually reach, and judged across the whole identity space rather than one attribute at a time.
Comments: 49 pages, 17 figures, 12 tables; Submitted & Accepted to ICAIF'2026
Subjects: Artificial Intelligence (cs.AI)
MSC classes: 68T50, 91G80, 62P20
ACM classes: J.4; I.2.7; K.4.2
Cite as: arXiv:2610.07798 [cs.AI]
  (or arXiv:2610.07798v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.07798

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Sankar B [view email]
[v1] Tue, 6 Oct 2026 05:46:28 UTC (2,484 KB)

来源:arXiv:cs.AI · arxiv.org