跳到正文
arXiv:cs.AI· Guanghao Wu, Zhuo Cai, Shoujin Wang·· 3 小时前

MemTrial:让 LLM 投资组合智能体学会何时信任记忆

MemTrial: Learning When to Trust Memory in LLM Portfolio Agents

AI 导读

针对 LLM 投资组合智能体因记忆信用绑定市场涨跌而跑输 1/N 等权组合的问题,MemTrial 通过分数析因设计生成同一决策的八个记忆组合草案,以 Banzhaf 值给每条经验记功,并用分层贝叶斯模型跨日期汇总、仅在预测过未见日期后才据此操作。

正文

View PDF HTML (experimental)

Abstract:Large language model (LLM) agents for portfolio management learn from experience: they credit each experience in their memory with the outcome of the decisions that used it. In financial markets, however, this outcome mostly reflects the market move shared by all decisions on that date, so the credit tracks the market rather than the experience, and these agents often do worse than simply holding the equal-weight (1/$N$) portfolio. We ask how an agent can credit an experience with what it changes, and answer it by putting memory on trial: drafts of the same decision with and without an experience face the same market, so the outcome they share cancels in their difference. Our agent, MemTrial, drafts each decision with eight combinations of its retrieved experiences, chosen by a fractional factorial design, and credits each experience with its Banzhaf value, the average of these differences. As each date occurs once and each draft is a noisy LLM sample, these credits are noisy and may not hold on new dates. MemTrial therefore pools them across dates and similar experiences with a hierarchical Bayesian model, acts on them only after they have predicted unseen dates, and otherwise stays anchored at a conservative reference such as 1/$N$. On four benchmarks, MemTrial not only benefits from experiences that matter (the best of 15 methods on a semi-synthetic benchmark with known experience quality) but also limits its losses when its values do not hold (at most 2.2\% below 1/$N$ on PortBench and InvestorBench, against 15--38\% for the best experience-learning agent). Averaged over five settings, it improves the utility of the best experience-learning agent by 21.2\%, and with eight LLMs it beats every LLM-based baseline on InvestorBench.
Subjects: Artificial Intelligence (cs.AI)
Cite as: arXiv:2610.11732 [cs.AI]
  (or arXiv:2610.11732v1 [cs.AI] for this version)
  https://doi.org/10.48550/arXiv.2610.11732

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Guanghao Wu [view email]
[v1] Thu, 8 Oct 2026 11:27:49 UTC (1,770 KB)

来源:arXiv:cs.AI · arxiv.org