arXiv:cs.LG· Siyu Wang, Yifan Wang, Yuecheng He·· 3 小时前AI 评分48
MintEval:LLM 是否真的实现了你要求的交易策略?自然语言转策略代码的行为等价性基准
MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code
AI 导读
MintEval 是一个衡量 LLM 将自然语言交易策略转为代码后行为是否等价的基准,v0 含 800 个 BTCUSDT 15 分钟数据任务。低成本模型平均 ActionMatch 最高仅 0.544,最多 0.087 的任务完全复现;Claude Opus 5.5 在 200 个分层任务上达 0.889、完全复现 0.575,但仍有 0.275 的任务静默失败。
正文
Abstract:Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed bar by bar on identical market data and frictions, and compared on their actions rather than on code similarity or profit: alpha is differenced away. MintEval v0 contains 800 tasks on BTCUSDT 15-minute data, stratified by an execution-measured state-span complexity tau that is decoupled from description length. Low-cost models reach a mean ActionMatch of at most 0.544 and reproduce at most 0.087 of tasks exactly; on a stratified subset of 200 tasks a frontier model (Claude Opus 5.5) reaches 0.889 and reproduces 0.575 exactly, yet still fails silently on 0.275 of tasks. Given a menu of building blocks, models identify the strategy almost perfectly, yet 79.2% of the implementations whose specification was read correctly diverge on more than 10% of active bars. The LLM judge of a recent strategy-generation benchmark, applied verbatim, accepts every one of these silent failures.
| Comments: | 5 pages, 3 figures, benchmark code and evaluation harness available at this https URL. Siyu Wang and Varstern Yifan Wang contributed equally, Yifig Wang is corresponding author |
| Subjects: | Software Engineering (cs.SE); Computation and Language (cs.CL); Machine Learning (cs.LG); Trading and Market Microstructure (q-fin.TR) |
| ACM classes: | D.2.5; I.2.7; J.4 |
| Cite as: | arXiv:2610.03080 [cs.SE] |
| (or arXiv:2610.03080v1 [cs.SE] for this version) | |
| https://doi.org/10.48550/arXiv.2610.03080 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yifan Wang [view email]
[v1]
Fri, 2 Oct 2026 10:01:26 UTC (210 KB)
来源:arXiv:cs.LG · arxiv.org