Reasoning Distilled Models
A collection of small language models fine-tuned on reasoning traces from multiple teacher models across 4 domains (math, science, if_chat, code).
Repository Structure
qwen2.5_1.5b/ # Qwen2.5-1.5B-Instruct base
βββ {parent}/
β βββ math/ # Single-domain (~18k)
β βββ science/
β βββ if_chat/
β βββ code/
β βββ code(1.5k)/ # Code-only, data-cardinality splits
β βββ code(4.5k)/
β βββ code(9k)/
β βββ mixed_domain(18k)/ # Mix of code/math/science/if_chat (18k)
β βββ mixed_domain(72k)/ # All 4 domains combined (72k)
βββ mixed_parent/
βββ code(54k)/ # Code from 3 teachers combined
βββ code(90k)/ # Code from 5 teachers combined
llama3.2_1b/ # Llama-3.2-1B-Instruct base
βββ {parent}/
β βββ code/ # Single-domain (code only)
β βββ mixed_domain(18k)/ # Mix of code/math/science/if_chat (18k)
β βββ mixed_domain(72k)/ # All 4 domains combined (72k)
qwen2.5_3b/ # Qwen2.5-3B-Instruct base
βββ {parent}/
β βββ code(18k)/ # Code-only traces (18k)
qwen2.5_7b/ # Qwen2.5-7B-Instruct base
βββ {parent}/
β βββ code(18k)/ # Code-only traces (18k)
β βββ code(1.5k)/ # Code-only, data-cardinality splits
β βββ code(4.5k)/
β βββ code(9k)/
qwen3_coder_models/ # Qwen2.5-Coder-1.5B-Instruct base
βββ {parent}/ # Code-only traces (18k), 10 parents
Teacher Models (Parents)
Exact teacher checkpoints, recovered from the trace-generation code:
| Parent | Teacher model |
|---|---|
deepseek_distill_llama3_70b |
deepseek-ai/DeepSeek-R1-Distill-Llama-70B |
deepseek_distill_qwen_25 |
deepseek-ai/DeepSeek-R1-Distill-Qwen-32B |
deepseekv3_2 |
DeepSeek-V3.2 β exact checkpoint unconfirmed |
exaone4_32b |
LGAI-EXAONE/EXAONE-4.0-32B |
glm |
zai-org/GLM-4.7-Flash |
gpt_oss_120b |
openai/gpt-oss-120b |
gpt_oss_20b |
openai/gpt-oss-20b |
olmo31 |
allenai/Olmo-3.1-32B-Think |
qwen3_235B |
Qwen/Qwen3-235B-A22B (bf16, thinking enabled) |
qwq |
Qwen/QwQ-32B |
Note that deepseek_distill_qwen_25 is the 32B R1 distill β the 25 refers to its
Qwen2.5 base, not a parameter count. Teacher traces were generated with vLLM in thinking
mode, reasoning kept inside <think>...</think>.
Training Details
- Epochs: 3
- Learning rate: 1e-4 with 10% warmup
- Max length: 32,768 tokens
- Precision: bf16
- Loss: Completion-only loss (assistant tokens only)
- Hardware: 8Γ H100 (torchrun; DeepSpeed ZeRO-2 for 1.5B, ZeRO-3 + CPU offload for 7B)
Qwen2.5-1.5B Models
All 10 parents unless noted.
| Variant | Models | Notes |
|---|---|---|
math / science / if_chat / code |
40 | 10 parents Γ 4 single domains, ~18k samples each |
code(1.5k) |
10 | Code-only data-cardinality split |
code(4.5k) |
10 | Code-only data-cardinality split |
code(9k) |
10 | Code-only data-cardinality split |
mixed_domain(18k) |
10 | Mix of all 4 domains (18k) |
mixed_domain(72k) |
10 | All 4 domains (72k) |
mixed_parent/code(54k) |
1 | Code traces mixed across 3 teachers |
mixed_parent/code(90k) |
1 | Code traces mixed across 5 teachers |
Data cardinality splits (code(1.5k) / code(4.5k) / code(9k)) hold the domain fixed at code and vary only the number of training traces, isolating the effect of data volume on lineage signal.
Mixed-parent models hold the domain fixed at code and blend traces from multiple teachers into one training set, so no single lineage dominates:
| Variant | Teachers | Traces |
|---|---|---|
mixed_parent/code(54k) |
gpt_oss_120b, deepseekv3_2, qwen3_235B |
3 Γ 18k = 54k |
mixed_parent/code(90k) |
the 3 above + exaone4_32b, qwq |
5 Γ 18k = 90k |
Llama-3.2-1B Models
- 1 single-domain model:
qwen3_235B/ code - 6
mixed_domain(18k)models - 6
mixed_domain(72k)models
Qwen2.5-3B Models
- 6
code(18k)models: 6 parents Γ code-only traces (18k)
Qwen2.5-7B Models
| Variant | Models | Parents |
|---|---|---|
code(18k) |
6 | all 6 |
code(1.5k) |
3 | deepseekv3_2, qwen3_235B, qwq |
code(4.5k) |
3 | deepseekv3_2, qwen3_235B, qwq |
code(9k) |
2 | qwen3_235B, qwq |
The 6 parents for 3B/7B models are: qwen3_235B, deepseekv3_2, gpt_oss_120b, olmo31, deepseek_distill_qwen_25, qwq
Qwen2.5-Coder-1.5B Models
qwen3_coder_models/{parent}/ β 10 models, one per parent, each SFT'd on that
teacher's 18k code-mix traces. Base model is
Qwen/Qwen2.5-Coder-1.5B-Instruct.
These use the same recipe as the qwen2.5_1.5b code models, so the only variable is the
base model β a code-specialised base instead of a general one. That makes the two sets
directly comparable for measuring how much of the lineage signal survives when the student
already has code pretraining.
| Base | Qwen/Qwen2.5-Coder-1.5B-Instruct |
| Parents | all 10 |
| Traces | 18,000 code-mix per parent |
| Steps | 423 (3 epochs, effective batch 128) |
The base swap is architecturally a no-op: Coder-1.5B-Instruct matches Qwen2.5-1.5B-Instruct
on every config field (28 layers, hidden 1536, vocab 151936, max_position_embeddings
32768) and the chat template is byte-identical.
Usage
from transformers import AutoModelForCausalLM, AutoTokenizer
# Qwen2.5-1.5B β single domain
model = AutoModelForCausalLM.from_pretrained(
"SupritiVijay/reasoning-distillation-models",
subfolder="qwen2.5_1.5b/gpt_oss_120b/math",
torch_dtype="auto",
device_map="auto",
)
# Qwen2.5-1.5B β code data-cardinality split
model = AutoModelForCausalLM.from_pretrained(
"SupritiVijay/reasoning-distillation-models",
subfolder="qwen2.5_1.5b/glm/code(4.5k)",
torch_dtype="auto",
device_map="auto",
)
# Qwen2.5-1.5B β mixed domain (72k)
model = AutoModelForCausalLM.from_pretrained(
"SupritiVijay/reasoning-distillation-models",
subfolder="qwen2.5_1.5b/qwen3_235B/mixed_domain(72k)",
torch_dtype="auto",
device_map="auto",
)
# Qwen2.5-1.5B β mixed parent (90k, code only)
model = AutoModelForCausalLM.from_pretrained(
"SupritiVijay/reasoning-distillation-models",
subfolder="qwen2.5_1.5b/mixed_parent/code(90k)",
torch_dtype="auto",
device_map="auto",
)
# Qwen2.5-Coder-1.5B β code (18k)
model = AutoModelForCausalLM.from_pretrained(
"SupritiVijay/reasoning-distillation-models",
subfolder="qwen3_coder_models/qwen3_235B",
torch_dtype="auto",
device_map="auto",
)
# Llama-3.2-1B β mixed domain (72k)
model = AutoModelForCausalLM.from_pretrained(
"SupritiVijay/reasoning-distillation-models",
subfolder="llama3.2_1b/qwen3_235B/mixed_domain(72k)",
torch_dtype="auto",
device_map="auto",
)
# Qwen2.5-3B β code (18k)
model = AutoModelForCausalLM.from_pretrained(
"SupritiVijay/reasoning-distillation-models",
subfolder="qwen2.5_3b/qwen3_235B/code(18k)",
torch_dtype="auto",
device_map="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
"SupritiVijay/reasoning-distillation-models",
subfolder="qwen2.5_3b/qwen3_235B/code(18k)",
)
Download a Single Model (CLI)
huggingface-cli download SupritiVijay/reasoning-distillation-models \
--include "qwen2.5_3b/gpt_oss_120b/code(18k)/*" \
--local-dir ./model