gguf-kernels / README.md
marcsun13's picture
marcsun13 HF Staff
GGUF kernels: dequantize + fused gemv over packed blocks
f79d690 verified
|
Raw
History Blame Contribute Delete
3.21 kB
metadata
license: mit
tags:
  - kernel
  - gguf
  - quantization

GGUF kernels

Compute directly on the packed blocks of a GGUF checkpoint.

A GGUF weight is stored as blocks of 32 or 256 values that share a scale. These kernels read that layout as it is, so a quantized model runs without ever materializing a dense copy of its weights — which is where the memory saving comes from. A 4B model at Q4_K_M stays at 3.1 GB instead of the 8.4 GB it would take unpacked.

Ported from llama.cpp's ggml-cuda. vendor/UPSTREAM pins the revision the sources come from; vendor.py refreshes them.

Usage

import torch
from kernels import get_kernel

gguf_kernels = get_kernel("marcsun13/gguf-kernels")

# `blocks` is a GGUF weight exactly as stored: (out_features, bytes_per_row) uint8
out = gguf_kernels.mul_mat_vec(blocks, x, ggml_type=12, out_features=4096)   # x @ blocks.T
dense = gguf_kernels.dequantize(blocks, ggml_type=12, rows=4096, cols=2560, dtype=torch.bfloat16)
mul_mat_vec(blocks, x, ggml_type, out_features) Fused dequantize-gemv, for x of at most MAX_GEMV_ROWS rows (decode). Returns f32 whatever x was: the kernel writes an f32 destination, and casting at the call site lets the cast fuse into the consumer.
dequantize(blocks, ggml_type, rows, cols, dtype) Blocks to values, in dtype directly. For more rows than the gemv handles (prefill), unpack and use an ordinary matmul.
MAX_GEMV_ROWS Upstream's MMVQ_MAX_BATCH_SIZE. Beyond this the gemv has no implementation.
GEMV_TYPES The ggml type ids mul_mat_vec implements. dequantize covers more, so check this before choosing the fused path.

Both ops carry register_fake implementations, so they trace under torch.compile instead of breaking the graph — which matters, since a break at every quantized linear costs more than the kernels save.

Backends

CUDA today. The ops are named for what they do rather than for a backend: the schema is declared once in torch-ext/torch_binding.cpp and each backend registers its own implementation of it, so adding Metal means adding a [kernel.*] section and a source directory — no change to the schema or to the Python API, and callers dispatch on the tensor's device as usual.

mmq (quantized gemm) and mmf (dense gemm) are deliberately not ported: both take a ggml_backend_cuda_context for their pool allocator, which would pull in the whole ggml backend. Above the gemv's row limit, unpacking and using cuBLAS measures at parity with llama.cpp anyway.

Building

nix run .#build-and-copy -L    # every variant, into build/
nix build .#ci                 # just this system's variant
pytest tests/                  # against a built variant on PYTHONPATH

flake.nix tracks the builder's default branch, so CI builds the current variant matrix. The build/ directory here was produced against the builder release matching the kernels client it was verified with (torch 2.11/2.12): both ops tested against gguf.quants.dequantize for Q4_K/Q5_K/Q6_K/ Q8_0 in f32/f16/bf16, the gemv against a dense matmul, and one fullgraph=True compile.