arXiv:cs.AI· Weixian Xu, Yanzhe Zhang, Zora Zhiruo Wang, Changyu Chen, Diyi Yang·· 6 小时前AI 评分49
Sherpa:教 LLM 自适应教学
Sherpa: Teaching LLMs to Teach Adaptively
AI 导读
研究团队提出多轮强化学习框架 Sherpa,通过让教师模型直接最大化多种模拟学生的学习效果来训练其自适应教学能力。经 Sherpa 训练的教师模型使学生成绩平均提升 20.5 个百分点,在 MathTutorBench 上的整体教学评分从 52.5% 升至 79.2%。人类研究显示,该教师在 79.6% 的成对比较中优于基座模型,代码与模型已公开。
正文
Abstract:Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
| Comments: | 32 pages, 6 figures. Code and model are available at this https URL |
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL) |
| Cite as: | arXiv:2610.08778 [cs.AI] |
| (or arXiv:2610.08778v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08778 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Weixian Xu [view email]
[v1]
Tue, 6 Oct 2026 17:58:18 UTC (585 KB)
来源:arXiv:cs.AI · arxiv.org