arXiv:cs.CL· Tong Xiao, Jingbo Zhu·· 3 小时前
大语言模型基础:预训练、生成、提示词、对齐、推理与推理能力
Foundations of Large Language Models
AI 导读
arXiv 上线《Foundations of Large Language Models》一书,聚焦大语言模型基础概念而非全部前沿技术,全书分为预训练、生成模型、提示词、对齐、推理、推理能力六个主要章节。该书面向高校学生、NLP 从业者及相关领域实践者,可作为大语言模型入门参考。
正文
Abstract:This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals, and practitioners in natural language processing and related fields, and can serve as a reference for anyone interested in large language models.
| Comments: | Minor corrections |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG) |
| Cite as: | arXiv:2501.09223 [cs.CL] |
| (or arXiv:2501.09223v4 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2501.09223 arXiv-issued DOI via DataCite |
Submission history
From: Tong Xiao [view email]
[v1]
Thu, 16 Jan 2025 01:03:56 UTC (361 KB)
[v2]
Sun, 15 Jun 2025 13:24:11 UTC (514 KB)
[v3]
Thu, 24 Sep 2026 15:06:31 UTC (942 KB)
[v4]
Thu, 8 Oct 2026 14:36:23 UTC (600 KB)
来源:arXiv:cs.CL · arxiv.org