How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf webAI-Official/granite-4.2-8B:
# Run inference directly in the terminal:
llama cli -hf webAI-Official/granite-4.2-8B:
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf webAI-Official/granite-4.2-8B:
# Run inference directly in the terminal:
llama cli -hf webAI-Official/granite-4.2-8B:
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf webAI-Official/granite-4.2-8B:
# Run inference directly in the terminal:
./llama-cli -hf webAI-Official/granite-4.2-8B:
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf webAI-Official/granite-4.2-8B:
# Run inference directly in the terminal:
./build/bin/llama-cli -hf webAI-Official/granite-4.2-8B:
Use Docker
docker model run hf.co/webAI-Official/granite-4.2-8B:
Quick Links

granite-4.2-8B GGUF

GGUF conversions of ibm-granite/granite-4.2-8b for llama.cpp at 16-bit, 8-bit and 4-bit precision.

Files

File Precision Size
granite-4.2-8b-BF16.gguf BF16 (16-bit, lossless from source) 17.6 GB
granite-4.2-8b-Q8_0.gguf Q8_0 (8-bit, 8.50 bits per weight) 9.3 GB
granite-4.2-8b-Q4_K_M.gguf Q4_K_M (4-bit, 4.86 bits per weight) 5.3 GB

Usage

# Chat server with Granite's embedded chat template (tool calling and thinking)
llama-server -hf webAI-Official/granite-4.2-8B:Q4_K_M --jinja

# Interactive chat
llama-cli -hf webAI-Official/granite-4.2-8B:Q8_0

Thinking is on by default in the chat template. Pass --chat-template-kwargs '{"enable_thinking": false}' to llama-server to turn it off.

How these files were made

  1. The source safetensors (bf16) were converted with llama.cpp b9770: convert_hf_to_gguf.py ibm-granite/granite-4.2-8b --outtype bf16
  2. Q8_0 and Q4_K_M were each quantized directly from the BF16 file with llama-quantize. No importance matrix was used.

Performance

Measured with llama-bench (llama.cpp b9770, Metal) on an Apple M5 Pro with 24 GB of memory:

File Prompt processing, 512 tokens (tok/s) Generation, 128 tokens (tok/s)
BF16 834 15.3
Q8_0 1080 29.1
Q4_K_M 1051 47.8

On a 24 GB machine, BF16 does not fit in GPU memory with llama.cpp's default context size. Use a smaller context, such as -c 2048, or choose Q8_0.

License

Apache 2.0, the same as the base model. See ibm-granite/granite-4.2-8b for the model's details, intended use and limitations.

Downloads last month
139
GGUF
Model size
9B params
Architecture
granite
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webAI-Official/granite-4.2-8B

Quantized
(47)
this model