DavidAU/Qwen3.5-9B-The-Defiant-Fable-Uncensored-Heretic-NEO-IMATRIX-MAX-MTP-GGUF Image-Text-to-Text • 9B • Updated 3 days ago • 379k • 317
view post Post 4342 We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. ⚡Qwen3.6-27B NVFP4 runs on 24GB VRAM.35B-A3B can hit 17,561 tok/s (B200).We also improved accuracy, tool calling, agent use, and looping.Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 See translation 1 reply · 🚀 15 15 🔥 11 11 🤗 1 1 + Reply
🚀 Qwen-MTP Collection ⚡ MTP (Multi Token Prediction) speculative decoding enables models like Qwen3.6 to have ~1.4-2.2x faster generation with no change in accuracy. • 9 items • Updated Jun 29 • 35
🍎 Qwopus3.6 Collection This collection features the advanced Qwopus3.6 series of multimodal large models, which are fine-tuned from the Qwen3.6 base models with a focus on e • 8 items • Updated Jul 1 • 77
💻 Qwopus-Coder Collection Reasoning-distilled coding models optimized for specialized domains like agentic workflows. • 10 items • Updated Jul 1 • 47