跳到正文
arXiv:cs.CL· Tong Xiao, Jingbo Zhu·· 3 小时前

大语言模型基础:预训练、生成、提示词、对齐、推理与推理能力

Foundations of Large Language Models

AI 导读

arXiv 上线《Foundations of Large Language Models》一书,聚焦大语言模型基础概念而非全部前沿技术,全书分为预训练、生成模型、提示词、对齐、推理、推理能力六个主要章节。该书面向高校学生、NLP 从业者及相关领域实践者,可作为大语言模型入门参考。

正文

View PDF HTML (experimental)

Abstract:This is a book about large language models. As indicated by the title, it primarily focuses on foundational concepts rather than comprehensive coverage of all cutting-edge technologies. The book is structured into six main chapters, each exploring a key area: pre-training, generative models, prompting, alignment, inference, and reasoning. It is intended for college students, professionals, and practitioners in natural language processing and related fields, and can serve as a reference for anyone interested in large language models.
Comments: Minor corrections
Subjects: Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Machine Learning (cs.LG)
Cite as: arXiv:2501.09223 [cs.CL]
  (or arXiv:2501.09223v4 [cs.CL] for this version)
  https://doi.org/10.48550/arXiv.2501.09223

arXiv-issued DOI via DataCite

Submission history

From: Tong Xiao [view email]
[v1] Thu, 16 Jan 2025 01:03:56 UTC (361 KB)
[v2] Sun, 15 Jun 2025 13:24:11 UTC (514 KB)
[v3] Thu, 24 Sep 2026 15:06:31 UTC (942 KB)
[v4] Thu, 8 Oct 2026 14:36:23 UTC (600 KB)

来源:arXiv:cs.CL · arxiv.org