DistilQwen
Proof-weighted distillation, Qwen3-30B to 1.7B/0.6B. Three teachers: Instruct, Thinking, Coder. The core method series. DOI 10.57967/hf/8165
Text Generation • 2B • Updated • 2.92k • • 2Note First in the DistilQwen chain. Foundation for all downstream models.
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B-SFT-GGUF
Text Generation • 2B • Updated • 2.99kNote Instruct teacher + SFT quantized. F16/Q4/Q5/Q8 available.
reaperdoesntknow/Qwen3-1.7B-Distilled-30B-A3B-SFT
2B • Updated • 153Note Source model for the SFT-GGUF quantizations.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B
Text Generation • 0.8B • Updated • 2.77k •Note Base for the Thinking-SFT pipeline at 0.6B.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT
Text Generation • 0.8B • Updated • 2.82k • • 2Note The smallest model with Thinking teacher signal. 0.8B params.
reaperdoesntknow/Qwen3-0.6B-Distilled-30B-A3B-Thinking-SFT-GGUF
Text Generation • 0.8B • Updated • 2.84kNote Edge deployment of extended-thinking at 0.6B. Apache 2.0.
reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT
Text Generation • 2B • Updated • 2.85k • 1Note Coder teacher produces uniquely structured distributions.
reaperdoesntknow/Qwen3-1.7B-Coder-Distilled-SFT-GGUF
Text Generation • 2B • Updated • 3.38k • 1Note Structured reasoning for edge deployment. Apache 2.0.
reaperdoesntknow/DistilQwen3-1.7B-uncensored
Text Generation • 2B • Updated • 2.22k •Note Foundation for research applications requiring unfiltered output.
reaperdoesntknow/TopologicalQwen
Text Generation • 2B • Updated • 3.41k • • 1Note TKD flagship. BV decomposition → jump detection → curriculum.
reaperdoesntknow/DiStil-Qwen3-1.7B-uncensored
2B • Updated • 154 • 1Note Named for Discrepancy Calculus influence on training signal.
reaperdoesntknow/Disctil-Qwen3-1.7B
Text Generation • 2B • Updated • 2.1k •Note Structural refinement via DISC operator before TKD stage.
reaperdoesntknow/DistilQwen3-1.7B-uncensored-GGUF
2B • Updated • 3.01k • 3Note Edge deployment for research. No alignment filtering. Apache 2.0.
reaperdoesntknow/Qwen3-1.7B-Thinking-Distil
Text Generation • 2B • Updated • 2.98k • • 2Note Extended deliberation from 30B-Thinking → 1.7B student.
reaperdoesntknow/LFM2.5-1.2B-Distilled-SFT
Text Generation • 1B • Updated • 2.36kNote Proves TKD works across architecture families, not just within Qwen.
reaperdoesntknow/Discrepancy_Calculus
UpdatedNote Continuous Thought Dynamics — mathematical backbone of DualMind.