GGUF kernels
Compute directly on the packed blocks of a GGUF checkpoint.
A GGUF weight is stored as blocks of 32 or 256 values that share a scale. These kernels read that layout as it is, so a quantized model runs without ever materializing a dense copy of its weights โ which is where the memory saving comes from. A 4B model at Q4_K_M stays at 3.1 GB instead of the 8.4 GB it would take unpacked.
Ported from llama.cpp's ggml-cuda. vendor/UPSTREAM
pins the revision the sources come from; vendor.py refreshes them.
Usage
import torch
from kernels import get_kernel
gguf_kernels = get_kernel("marcsun13/gguf-kernels")
# `blocks` is a GGUF weight exactly as stored: (out_features, bytes_per_row) uint8
out = gguf_kernels.mul_mat_vec(blocks, x, ggml_type=12, out_features=4096) # x @ blocks.T
dense = gguf_kernels.dequantize(blocks, ggml_type=12, rows=4096, cols=2560, dtype=torch.bfloat16)
mul_mat_vec(blocks, x, ggml_type, out_features) |
Fused dequantize-gemv, for x of at most MAX_GEMV_ROWS rows (decode). Returns f32 whatever x was: the kernel writes an f32 destination, and casting at the call site lets the cast fuse into the consumer. |
dequantize(blocks, ggml_type, rows, cols, dtype) |
Blocks to values, in dtype directly. For more rows than the gemv handles (prefill), unpack and use an ordinary matmul. |
MAX_GEMV_ROWS |
Upstream's MMVQ_MAX_BATCH_SIZE. Beyond this the gemv has no implementation. |
GEMV_TYPES |
The ggml type ids mul_mat_vec implements. dequantize covers more, so check this before choosing the fused path. |
Both ops carry register_fake implementations, so they trace under torch.compile instead of
breaking the graph โ which matters, since a break at every quantized linear costs more than the
kernels save.
Backends
CUDA today. The ops are named for what they do rather than for a backend: the schema is declared
once in torch-ext/torch_binding.cpp and each backend registers its own implementation of it, so
adding Metal means adding a [kernel.*] section and a source directory โ no change to the schema or
to the Python API, and callers dispatch on the tensor's device as usual.
mmq (quantized gemm) and mmf (dense gemm) are deliberately not ported: both take a
ggml_backend_cuda_context for their pool allocator, which would pull in the whole ggml backend.
Above the gemv's row limit, unpacking and using cuBLAS measures at parity with llama.cpp anyway.
Building
nix run .#build-and-copy -L # every variant, into build/
nix build .#ci # just this system's variant
pytest tests/ # against a built variant on PYTHONPATH
flake.nix tracks the builder's default branch, so CI builds the current variant matrix. The
build/ directory here was produced against the builder release matching the kernels client it was
verified with (torch 2.11/2.12): both ops tested against gguf.quants.dequantize for Q4_K/Q5_K/Q6_K/
Q8_0 in f32/f16/bf16, the gemv against a dense matmul, and one fullgraph=True compile.