Training Specialist Models without Reasoning Trajectories for Domain Expert Distillation
Abstract
Specialist training without explicit reasoning supervision implicitly selects latent reasoning trajectories that govern distilled students' specialization and generalization trade-offs.
Specialist distillation effectively transfers domain expertise to student models via teacher-generated reasoning trajectories. However, when these specialists are trained solely on question--answer pairs without explicit reasoning supervision, what governs the trajectories they generate? In this work, we show that specialist optimization implicitly selects from this latent trajectory space. To isolate and observe this latent distribution, we leverage student distillation not as a downstream goal, but as an agnostic probe---since students inherit no parameterization or optimization constraints from the specialist, inheriting only the sampled trajectories themselves. Through this probe, our empirical analysis unveils a tight governing relationship: across 27 specialist--student pairings, their specialization--generalization profiles correlate exceptionally strongly. Crucially, explicitly controlling the specialist's distributional drift systematically shifts both the teacher and its distilled student along a controllable trade-off between domain precision and general-capability retention. Across chemistry, physics, and multilingual settings, distilled students systematically reflect these specialist-induced profiles, even across divergent model families. Our findings establish a new view of specialist training: when gold reasoning is absent, tuning choices directly control the latent supervision passed to downstream models.
Community
A specialist is trained on answers only, with no reasoning supervision.
So what determines the reasoning trajectories it later generates to teach a student?
We find that specialist training itself implicitly selects these latent trajectories. Across 27 specialist–student pairs, students inherit their teachers’ specialization–generalization profiles—even across different model families.
And by controlling how the specialist is trained, we can systematically control what kind of reasoning supervision gets passed downstream.
Get this paper in your agent:
hf papers read 2609.13770 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper