arXiv:cs.LG(机器学习,全量分类)· Dong Shu, Yanguang Liu, Huopu Zhang, Saisai Hu, Haiyan Zhao, Hekun Huang, Mengnan Du·· 5 小时前AI 评分40
MM-FinEval:面向真实金融预测的多任务多模态基准
MM-FinEval: A Multi-Task Multimodal Benchmark for Real-World Financial Forecasting
AI 导读
研究者提出 MM-FinEval,一个覆盖 2019 至 2022 年、包含 2,045 场 S&P 500 财报电话会议和 12 项金融任务标签的多任务多模态基准,每场输入含文本记录、演示幻灯片和完整音频三种模态。
正文
Abstract:Financial forecasting from earnings conference calls requires models to reason over complex corporate disclosures, market expectations, and subtle communication signals. However, existing financial benchmarks are often limited to unimodal inputs or single-task settings, making it difficult to evaluate whether multimodal large language models (LLMs) can support real-world financial analysis. In this paper, we introduce MM-FinEval, a novel benchmark designed to evaluate multimodal LLMs across multiple financial tasks. MM-FinEval spans a diverse timeline from 2019 to 2022. The entire proposed dataset contains 2,045 S\&P 500 conference earning calls as inputs and 12 financial task labels as outputs. Each input contains three modalities: a word-to-word text transcript of the earning call, the corresponding presentation slides used during the call, and the entire audio recording. To establish a rigorous evaluation framework, we analyze 19 baseline models across three distinct model categories: Image-Text, Audio-Text, and Any-to-Any configurations. We observe that small-size Any-to-Any models processing all three modalities achieve strong performance, even when compared against larger proprietary models restricted to two-modality inputs. This indicates that our tri-modal dataset design introduces useful, non-redundant information. These results validate that text, audio, and visual data serve as important, complementary signals that mimic the decision-making process of expert human analysts.
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2609.38523 [cs.LG] |
| (or arXiv:2609.38523v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2609.38523 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Dong Shu [view email]
[v1]
Tue, 29 Sep 2026 20:40:42 UTC (6,827 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org