arXiv:cs.LG(机器学习,全量分类)· Geon-Woo Kim, Joon Ha Kim, Daehyeok Kim·· 13 小时前AI 评分40
Leto:在存活硬件上实现 LLM 训练的快速原地恢复
Leto: Fast In-Place Recovery for LLM Training on Surviving Hardware
AI 导读
Leto 是一个容错训练系统,利用存活硬件实现 LLM 训练的原地恢复,在 6 卡和 72 卡 NVIDIA A100 集群上恢复速度比最佳 checkpointing 基线快 3.6–6.5 倍,有效训练时间最多提升 13.7 个百分点。
正文
Abstract:Hardware-operable failures (HOFs) interrupt large language model (LLM) training but permit recovery on the same hardware without reset, repair, or replacement. Existing recovery systems nevertheless reload checkpoints, recompute lost progress, and rebuild process state, idling GPUs that could otherwise continue training.
We present Leto, a fault-tolerant training system that leverages surviving hardware to enable efficient in-place recovery. Our key insight is that the state needed to resume training can be retained or prepared outside the active training process while remaining on the same hardware. Leto retains the working model state and the reusable process state, and preinitializes the remaining state in a shadow trainer. We devise two-tier erasure protection and chunk-level transactional updates to keep the retained model state recoverable and consistent, and reclaim the shadow state when active training needs its GPU memory. Evaluation on 6- and 72-GPU NVIDIA A100 clusters shows that Leto recovers 3.6--6.5$\times$ faster than the best-performing checkpointing baselines and improves productive training time by up to 13.7 percentage points. Large-scale simulation shows over 95% productive training time on a 131,072-GPU cluster.
| Subjects: | Distributed, Parallel, and Cluster Computing (cs.DC); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.00687 [cs.DC] |
| (or arXiv:2610.00687v1 [cs.DC] for this version) | |
| https://doi.org/10.48550/arXiv.2610.00687 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Geon-Woo Kim [view email]
[v1]
Wed, 30 Sep 2026 20:29:57 UTC (2,792 KB)
来源:arXiv:cs.LG(机器学习,全量分类) · arxiv.org