How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
Use Docker
docker model run hf.co/thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
Quick Links

โšก Quantized with QuantizeLab

This model was quantized to GGUF format using QuantizeLab, the fastest SaaS to quantize models under 32B parameters. Join hundreds of developers downloading our optimized quants!

  • Fast Conversion: Direct upload to Hugging Face.
  • Optimized Sizes: Support for all major GGUF bit-rates.
  • Hardware Free: No local GPU required for quantization.

๐Ÿ‘‰ Try QuantizeLab Now

Qwen2.5-32B-Instruct-GGUF (Q4_K_M)

Q4_K_M GGUF quantization of Qwen/Qwen2.5-32B-Instruct, produced with llama.cpp by Quantizelab.dev.

File model-Q4_K_M.gguf
Quantization Q4_K_M
Size on disk 19.85 GB
Base model Qwen/Qwen2.5-32B-Instruct

Run it (full GPU offload)

GGUF defaults to CPU. To get GPU speed you must offload every layer โ€” a single layer left on the CPU takes generation from ~25 tok/s to ~3 tok/s. -ngl 999 simply means "offload all of them".

# llama.cpp
llama-cli -hf thecodehaider/Qwen2.5-32B-Instruct-GGUF:model-Q4_K_M.gguf -ngl 999 -c 4096 -p "Hello"

# local file
llama-cli -m model-Q4_K_M.gguf -ngl 999 -c 4096 -cnv

# OpenAI-compatible server
llama-server -m model-Q4_K_M.gguf -ngl 999 -c 4096 --port 8080
# Ollama
ollama run hf.co/thecodehaider/Qwen2.5-32B-Instruct-GGUF:Q4_K_M
# llama-cpp-python
from llama_cpp import Llama
llm = Llama(model_path="model-Q4_K_M.gguf", n_gpu_layers=-1, n_ctx=4096)
print(llm("Hello", max_tokens=128)["choices"][0]["text"])

Will it fit your GPU?

Weights plus ~1.2 GB of KV-cache and compute overhead at a 4k context.

GPU VRAM Fits fully offloaded? Headroom for context
NVIDIA T4 / RTX 4060 16 GB No offload partially (-ngl lower) or use CPU
RTX 3090 / 4090 / A10 24 GB Tight ~2.9 GB (cap context ~2048)
A100 40GB 40 GB Yes ~18.9 GB (4k+ context)

If a row says No, lower -ngl until it fits, or run on CPU (GGUF works either way โ€” it is just slower).

Notes

  • Q4_K_M is the recommended balance of size and quality; Q8_0 and above will not fully offload to a 16 GB card for models past ~8B.
  • Reduce -c (context) first when you hit out-of-memory: the KV cache grows linearly with context length.
Downloads last month
401
GGUF
Model size
33B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for thecodehaider/Qwen2.5-32B-Instruct-GGUF

Base model

Qwen/Qwen2.5-32B
Quantized
(154)
this model