Abstract
Large language models (LLMs) have become increasingly capable problem solvers, but being able to solve a problem is not the same as being able to teach it. Existing approaches to training LLMs as teachers rely on demonstrations, preference data, or predefined pedagogical criteria that specify what good teaching looks like. However, these signals are often not grounded in individual student learning outcomes, where effective teaching strategies can vary substantially across learners. To address this, we introduce Sherpa, a multi-turn reinforcement learning framework that instantiates multiple student archetypes with LLMs conditioned on distinct learning preferences and trains a teacher model to adapt its instruction by directly maximizing their learning outcomes. Teacher LLMs trained with Sherpa improve instructed students' performance across all archetypes by an average of 20.5 percentage points. Under MathTutorBench's evaluation, Sherpa raises the overall pedagogy score from 52.5% to 79.2%, indicating better teaching responses. Our human studies show that the trained teacher is preferred over the base model in 79.6% of pairwise comparisons. Together, Sherpa trains LLM teachers to adapt to diverse simulated students and become better aligned with human teachers, paving the road towards AI tutors teaching real students.
Community
Sherpa trains LLMs to become adaptive teachers by optimizing directly for student learning outcomes rather than predefined notions of good teaching. Using multi-turn RL with diverse simulated student archetypes, the teacher learns to infer different students’ needs and adapt its teaching strategy accordingly. Sherpa improves student performance by 20.5 percentage points on average, raises MathTutorBench pedagogy scores from 52.5% to 79.2%, and is preferred over the base model by human teachers in 79.6% of comparisons.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- The Assistance Dilemma: Learning to Teach via Multi-Turn Reinforcement Learning (2026)
- How Should Teachers Be Prepared? RL on Student-Induced States for On-Policy Distillation (2026)
- OPSRD: On-Policy Self-Role Distillation (2026)
- StudentSim: Training LLM-based Student Simulators (2026)
- DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation (2026)
- SIPO: Unifying Reinforcement Learning with On-Policy Self-Distillation (2026)
- DualOPSD: Adaptive Privileged Teachers for On-Policy Self-Distillation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2610.08778 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper