thecodehaider's picture
Update README.md
2cd9fda verified
|
Raw
History Blame Contribute Delete
2.65 kB
---
library_name: gguf
base_model: Qwen/Qwen2.5-32B-Instruct
tags:
- gguf
- llama.cpp
- quantized
- q4_k_m
---
## âš¡ Quantized with QuantizeLab
This model was quantized to GGUF format using [QuantizeLab](https://quantizelab.dev), the fastest SaaS to quantize models under 32B parameters. Join hundreds of developers downloading our optimized quants!
* **Fast Conversion:** Direct upload to Hugging Face.
* **Optimized Sizes:** Support for all major GGUF bit-rates.
* **Hardware Free:** No local GPU required for quantization.
👉 **[Try QuantizeLab Now](https://quantizelab.dev)**
# Qwen2.5-32B-Instruct-GGUF (Q4_K_M)
`Q4_K_M` GGUF quantization of [Qwen/Qwen2.5-32B-Instruct](https://huggingface.co/Qwen/Qwen2.5-32B-Instruct),
produced with [llama.cpp](https://github.com/ggerganov/llama.cpp) by
[Quantizelab.dev](https://quantizelab.dev).
| | |
| --- | --- |
| File | `model-Q4_K_M.gguf` |
| Quantization | `Q4_K_M` |
| Size on disk | 19.85 GB |
| Base model | [Qwen/Qwen2.5-32B-Instruct](https://huggingface.co/Qwen/Qwen2.5-32B-Instruct) |
## Run it (full GPU offload)
GGUF defaults to CPU. To get GPU speed you must offload **every** layer — a
single layer left on the CPU takes generation from ~25 tok/s to ~3 tok/s.
`-ngl 999` simply means "offload all of them".
```bash
# llama.cpp
llama-cli -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:model-Q4_K_M.gguf -ngl 999 -c 4096 -p "Hello"
# local file
llama-cli -m model-Q4_K_M.gguf -ngl 999 -c 4096 -cnv
# OpenAI-compatible server
llama-server -m model-Q4_K_M.gguf -ngl 999 -c 4096 --port 8080
```
```bash
# Ollama
ollama run hf.co/thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
```
```python
# llama-cpp-python
from llama_cpp import Llama
llm = Llama(model_path="model-Q4_K_M.gguf", n_gpu_layers=-1, n_ctx=4096)
print(llm("Hello", max_tokens=128)["choices"][0]["text"])
```
## Will it fit your GPU?
Weights plus ~1.2 GB of KV-cache and compute overhead at a 4k context.
| GPU | VRAM | Fits fully offloaded? | Headroom for context |
| --- | ---- | --------------------- | -------------------- |
| NVIDIA T4 / RTX 4060 | 16 GB | No | offload partially (`-ngl` lower) or use CPU |
| RTX 3090 / 4090 / A10 | 24 GB | Tight | ~2.9 GB (cap context ~2048) |
| A100 40GB | 40 GB | Yes | ~18.9 GB (4k+ context) |
If a row says **No**, lower `-ngl` until it fits, or run on CPU (GGUF works
either way — it is just slower).
## Notes
- `Q4_K_M` is the recommended balance of size and quality; `Q8_0` and above
will not fully offload to a 16 GB card for models past ~8B.
- Reduce `-c` (context) first when you hit out-of-memory: the KV cache grows
linearly with context length.