arXiv:cs.LG· Nickil Maveli, Antonio Vergari, Shay B. Cohen·· 3 小时前AI 评分38
LLM 是在合成还是回忆?AlgoREval 基准评估算法代码检索能力
Are you Synthesizing or Recalling? Evaluating LLMs on Algorithmic Code Retrieval
AI 导读
研究者提出 AlgoREval 基准,用 599 道题覆盖 77 种经典算法、14 个领域、7 种编程语言和 4 种图输入表示,在零样本设置下评估 15 个 7B–34B 参数模型的参数化代码检索能力。
正文
Abstract:Large language models (LLMs) have demonstrated strong performance in code generation, where success depends on both recalling relevant algorithmic knowledge and reasoning about how to apply it. However, existing LLM pipelines are opaque, with no explicit separation between these two components. We argue that for well-known algorithms whose canonical implementations are widely accessible in pretraining corpora, code generation is better measured as \textit{parametric code retrieval}: reproducing a named algorithm from internalised knowledge rather than synthesizing a novel one. We introduce AlgoREval, a benchmark of 599 problems spanning classical 77 algorithms across 14 domains, 7 programming languages, and 4 graph-input representations to evaluate this capability in isolation, and assess 15 models (7B--34B parameters) in a zero-shot setting. We find substantial variation in retrieval accuracy across languages and input representations, even for widely documented algorithms and show that prompt augmentation with retrieved code snippets or structured algorithmic hints improve accuracy on complex algorithms, while SFT achieves broader language gains and GRPO achieves larger per-language gains on specific languages. Together, our results establish parametric code retrieval as a distinct, measurable capability and caution against deploying AI-generated algorithmic code without systematic validation.\footnote{Code and dataset are available at this https URL
| Comments: | 30 pages (preprint) |
| Subjects: | Machine Learning (cs.LG); Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Programming Languages (cs.PL) |
| Cite as: | arXiv:2610.02438 [cs.LG] |
| (or arXiv:2610.02438v1 [cs.LG] for this version) | |
| https://doi.org/10.48550/arXiv.2610.02438 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Nickil Maveli [view email]
[v1]
Thu, 1 Oct 2026 20:04:14 UTC (1,330 KB)
来源:arXiv:cs.LG · arxiv.org