--- library_name: gguf base_model: Qwen/Qwen2.5-32B-Instruct tags: - gguf - llama.cpp - quantized - q4_k_m --- ## ⚡ Quantized with QuantizeLab This model was quantized to GGUF format using [QuantizeLab](https://quantizelab.dev), the fastest SaaS to quantize models under 32B parameters. Join hundreds of developers downloading our optimized quants! * **Fast Conversion:** Direct upload to Hugging Face. * **Optimized Sizes:** Support for all major GGUF bit-rates. * **Hardware Free:** No local GPU required for quantization. 👉 **[Try QuantizeLab Now](https://quantizelab.dev)** # Qwen2.5-32B-Instruct-GGUF (Q4_K_M) `Q4_K_M` GGUF quantization of [Qwen/Qwen2.5-32B-Instruct](https://huggingface.co/Qwen/Qwen2.5-32B-Instruct), produced with [llama.cpp](https://github.com/ggerganov/llama.cpp) by [Quantizelab.dev](https://quantizelab.dev). | | | | --- | --- | | File | `model-Q4_K_M.gguf` | | Quantization | `Q4_K_M` | | Size on disk | 19.85 GB | | Base model | [Qwen/Qwen2.5-32B-Instruct](https://huggingface.co/Qwen/Qwen2.5-32B-Instruct) | ## Run it (full GPU offload) GGUF defaults to CPU. To get GPU speed you must offload **every** layer — a single layer left on the CPU takes generation from ~25 tok/s to ~3 tok/s. `-ngl 999` simply means "offload all of them". ```bash # llama.cpp llama-cli -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:model-Q4_K_M.gguf -ngl 999 -c 4096 -p "Hello" # local file llama-cli -m model-Q4_K_M.gguf -ngl 999 -c 4096 -cnv # OpenAI-compatible server llama-server -m model-Q4_K_M.gguf -ngl 999 -c 4096 --port 8080 ``` ```bash # Ollama ollama run hf.co/thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M ``` ```python # llama-cpp-python from llama_cpp import Llama llm = Llama(model_path="model-Q4_K_M.gguf", n_gpu_layers=-1, n_ctx=4096) print(llm("Hello", max_tokens=128)["choices"][0]["text"]) ``` ## Will it fit your GPU? Weights plus ~1.2 GB of KV-cache and compute overhead at a 4k context. | GPU | VRAM | Fits fully offloaded? | Headroom for context | | --- | ---- | --------------------- | -------------------- | | NVIDIA T4 / RTX 4060 | 16 GB | No | offload partially (`-ngl` lower) or use CPU | | RTX 3090 / 4090 / A10 | 24 GB | Tight | ~2.9 GB (cap context ~2048) | | A100 40GB | 40 GB | Yes | ~18.9 GB (4k+ context) | If a row says **No**, lower `-ngl` until it fits, or run on CPU (GGUF works either way — it is just slower). ## Notes - `Q4_K_M` is the recommended balance of size and quality; `Q8_0` and above will not fully offload to a 16 GB card for models past ~8B. - Reduce `-c` (context) first when you hit out-of-memory: the KV cache grows linearly with context length.