llama.cpp c1d0e7a, CUDA 12.8.1, Linux x86-64
This is a CUDA build of llama.cpp at commit
c1d0e7a004015f23bc0233470b747b596f29b264 (release b10621), for Linux x86-64. llama.cpp
publishes no Linux CUDA archive for that commit, so this one is built from source by a pinned,
reproducible recipe, which is included here.
| File | Bytes | SHA-256 |
|---|---|---|
llama-c1d0e7a-cuda12.8.1-linux-x64.tar.gz |
191,811,056 | 7fd27fca6917b1f5aeae38eef9815eed932cf1f7c490227e0eb15612cb238a21 |
Files here are never overwritten. A new build gets a new name, so a pinned URL and its hash stay true together.
What it contains
llama-server with its libraries, built for CUDA architectures 50, 52, 60, 61, 70, 75, 80, 86, 89,
90, 100 and 120, with the web UI off. BUILD-INFO.txt records the source commit, the build image
digest, the compiler, the Ubuntu snapshot, the package versions and the SHA-256 of every recipe
file.
NVIDIA's libraries are not included
The server also needs NVIDIA's CUDA runtime and cuBLAS, which are fetched from NVIDIA and placed
beside llama-server:
| Archive | SHA-256 |
|---|---|
cuda_cudart-linux-x86_64-12.8.90-archive.tar.xz |
8d566b5fe745c46842dc16945cf36686227536decd2302c372be86da37faca68 |
libcublas-linux-x86_64-12.8.4.1-archive.tar.xz |
21718957c2cf000bacd69d36c95708a2319199e39e056f8b4f0f68e3b9f323bb |
Both are from https://developer.download.nvidia.com/compute/cuda/redist/. Copy libcudart.so*,
libcublas.so* and libcublasLt.so* from their lib/ folders. The GPU driver comes from the
system.
How it was built
Run recipe/scripts/build_runtime.sh. It builds in the pinned nvidia/cuda:12.8.1-devel-ubuntu22.04
image, with packages from a fixed Ubuntu snapshot. Every source file gets a fixed time. The
compiler runs with address randomization off and a fixed process ID. Docker's default seccomp
profile is used, plus one personality rule. Before packaging, the build's CMake cache is checked
against the expected configuration. recipe/scripts/check_runtime_reproducible.sh builds twice and
requires the same archive. It did here: two independent builds gave identical bytes.
Qualification
It was checked on an NVIDIA RTX 3090, with Qwen3.5-9B Q4_K_M (SHA-256 cd76ec20…2a13), a 32,768-token
context, flash attention and thinking off. The test set was 33 labeled invoices, run twice:
- 33 of 33 documents completed in each run, at a mean of 13.1 and 13.4 seconds per document, with the slowest at 41.1 s;
- 271 of 273 line items were exact in each run;
- the two runs' answers were identical on 33 of 33;
- the answers were identical, on 33 of 33, to those of the earlier build this one replaces;
- peak graphics memory for the server was 6,393 MiB.
License
llama.cpp is MIT-licensed; its license is in LICENSE. The build recipe in recipe/ is also
offered under the MIT license.