跳到正文
arXiv:cs.CL· Hope McGovern, Anna Dolganov, Samuel Belkadi, Guillaume Kunsch, Dimitris Vlitas, David A. Smith·· 4 小时前AI 评分47

Apollo Restore:面向古希腊文献补缺的 240 亿参数基础大模型

Apollo Restore: A Foundation LLM for Historical Greek Optimized for Fill-in-the-Middle Restoration of Ancient Greek Texts

AI 导读

奥地利科学院主导的 Decoding Antiquity 项目发布 Apollo Restore,一个基于 Mistral Small 微调的 240 亿参数大模型,采用 fill-in-the-middle 目标修复古希腊残卷的物理缺失。

正文

View PDF HTML (experimental)

Abstract:We present Apollo Restore, a 24-billion-parameter large language model for restoring lacunae---physical gaps---in fragmentary Ancient Greek texts. Fine-tuned from Mistral Small with a fill-in-the-middle objective, Apollo Restore reconstructs missing spans without requiring oracle knowledge of their length. To our knowledge, it is the first large-scale decoder model for historical Greek, and the first for any ancient Mediterranean language. Evaluated as in prior work, on short gaps of up to ten characters, Apollo Restore places the correct restoration among its top twenty candidates for 80.6%/54.6%/61.0% of documentary-papyrus, literary-papyrus, and stone-inscription lacunae, exceeding the strongest published models by $1.6\times$/$2.6\times$/$1.4\times$. Prior evaluation protocols, however, inflate scores through a bias toward trivially short gaps; under a length-balanced metric Apollo Restore's advantage over the strongest published models grows to $2.3\times$/$3.5\times$/$1.6\times$ and degrades gracefully, even given incorrect length hints. In a blind study, 20 expert papyrologists, epigraphists, and philologists strongly preferred Apollo Restore to the strongest baseline and judged its performance at least as good as human restorations in 77% of cases. Apollo Restore also improves the published reading of PHerc. 1667---a papyrus roll carbonised in the eruption of Vesuvius in 79 CE and digitally unrolled and edited after Apollo Restore's training data was compiled. Apollo Restore is an output of the Decoding Antiquity initiative to build specialized LLMs for historical languages and manuscripts, led by the Austrian Academy of Sciences.
Comments: 16 pages, 6 figures
Subjects: Computation and Language (cs.CL)
Cite as: arXiv:2609.22455 [cs.CL]
  (or arXiv:2609.22455v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2609.22455

arXiv-issued DOI via DataCite

Submission history

From: David Smith [view email]
[v1] Fri, 18 Sep 2026 18:12:22 UTC (2,011 KB)
[v2] Tue, 22 Sep 2026 02:00:57 UTC (2,011 KB)
[v3] Thu, 24 Sep 2026 19:01:27 UTC (2,011 KB)
[v4] Wed, 7 Oct 2026 03:08:11 UTC (2,011 KB)

来源:arXiv:cs.CL · arxiv.org