File size: 3,211 Bytes
2c20074
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
f79d690
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
---
license: mit
tags:
  - kernel
  - gguf
  - quantization
---

# GGUF kernels

Compute directly on the packed blocks of a GGUF checkpoint.

A GGUF weight is stored as blocks of 32 or 256 values that share a scale. These kernels read that
layout as it is, so a quantized model runs without ever materializing a dense copy of its weights —
which is where the memory saving comes from. A 4B model at Q4_K_M stays at 3.1 GB instead of the
8.4 GB it would take unpacked.

Ported from [llama.cpp](https://github.com/ggml-org/llama.cpp)'s `ggml-cuda`. `vendor/UPSTREAM`
pins the revision the sources come from; `vendor.py` refreshes them.

## Usage

```python
import torch
from kernels import get_kernel

gguf_kernels = get_kernel("marcsun13/gguf-kernels")

# `blocks` is a GGUF weight exactly as stored: (out_features, bytes_per_row) uint8
out = gguf_kernels.mul_mat_vec(blocks, x, ggml_type=12, out_features=4096)   # x @ blocks.T
dense = gguf_kernels.dequantize(blocks, ggml_type=12, rows=4096, cols=2560, dtype=torch.bfloat16)
```

| | |
| --- | --- |
| `mul_mat_vec(blocks, x, ggml_type, out_features)` | Fused dequantize-gemv, for `x` of at most `MAX_GEMV_ROWS` rows (decode). Returns f32 whatever `x` was: the kernel writes an f32 destination, and casting at the call site lets the cast fuse into the consumer. |
| `dequantize(blocks, ggml_type, rows, cols, dtype)` | Blocks to values, in `dtype` directly. For more rows than the gemv handles (prefill), unpack and use an ordinary matmul. |
| `MAX_GEMV_ROWS` | Upstream's `MMVQ_MAX_BATCH_SIZE`. Beyond this the gemv has no implementation. |
| `GEMV_TYPES` | The ggml type ids `mul_mat_vec` implements. `dequantize` covers more, so check this before choosing the fused path. |

Both ops carry `register_fake` implementations, so they trace under `torch.compile` instead of
breaking the graph — which matters, since a break at every quantized linear costs more than the
kernels save.

## Backends

CUDA today. The ops are named for what they do rather than for a backend: the schema is declared
once in `torch-ext/torch_binding.cpp` and each backend registers its own implementation of it, so
adding Metal means adding a `[kernel.*]` section and a source directory — no change to the schema or
to the Python API, and callers dispatch on the tensor's device as usual.

`mmq` (quantized gemm) and `mmf` (dense gemm) are deliberately not ported: both take a
`ggml_backend_cuda_context` for their pool allocator, which would pull in the whole ggml backend.
Above the gemv's row limit, unpacking and using cuBLAS measures at parity with llama.cpp anyway.

## Building

```bash
nix run .#build-and-copy -L    # every variant, into build/
nix build .#ci                 # just this system's variant
pytest tests/                  # against a built variant on PYTHONPATH
```

`flake.nix` tracks the builder's default branch, so CI builds the current variant matrix. The
`build/` directory here was produced against the builder release matching the `kernels` client it was
verified with (torch 2.11/2.12): both ops tested against `gguf.quants.dequantize` for Q4_K/Q5_K/Q6_K/
Q8_0 in f32/f16/bf16, the gemv against a dense matmul, and one `fullgraph=True` compile.