Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss Paper • 2608.03796 • Published 4 days ago • 4
MultiverseComputingCAI/Qwen3-Next-80B-A3B-Thinking-Uncensored Text Generation • 80B • Updated Jun 2 • 151 • 16