llama.cpp c1d0e7a, CUDA 12.8.1, Linux x86-64

This is a CUDA build of llama.cpp at commit c1d0e7a004015f23bc0233470b747b596f29b264 (release b10621), for Linux x86-64. llama.cpp publishes no Linux CUDA archive for that commit, so this one is built from source by a pinned, reproducible recipe, which is included here.

File Bytes SHA-256
llama-c1d0e7a-cuda12.8.1-linux-x64.tar.gz 191,811,056 7fd27fca6917b1f5aeae38eef9815eed932cf1f7c490227e0eb15612cb238a21

Files here are never overwritten. A new build gets a new name, so a pinned URL and its hash stay true together.

What it contains

llama-server with its libraries, built for CUDA architectures 50, 52, 60, 61, 70, 75, 80, 86, 89, 90, 100 and 120, with the web UI off. BUILD-INFO.txt records the source commit, the build image digest, the compiler, the Ubuntu snapshot, the package versions and the SHA-256 of every recipe file.

NVIDIA's libraries are not included

The server also needs NVIDIA's CUDA runtime and cuBLAS, which are fetched from NVIDIA and placed beside llama-server:

Archive SHA-256
cuda_cudart-linux-x86_64-12.8.90-archive.tar.xz 8d566b5fe745c46842dc16945cf36686227536decd2302c372be86da37faca68
libcublas-linux-x86_64-12.8.4.1-archive.tar.xz 21718957c2cf000bacd69d36c95708a2319199e39e056f8b4f0f68e3b9f323bb

Both are from https://developer.download.nvidia.com/compute/cuda/redist/. Copy libcudart.so*, libcublas.so* and libcublasLt.so* from their lib/ folders. The GPU driver comes from the system.

How it was built

Run recipe/scripts/build_runtime.sh. It builds in the pinned nvidia/cuda:12.8.1-devel-ubuntu22.04 image, with packages from a fixed Ubuntu snapshot. Every source file gets a fixed time. The compiler runs with address randomization off and a fixed process ID. Docker's default seccomp profile is used, plus one personality rule. Before packaging, the build's CMake cache is checked against the expected configuration. recipe/scripts/check_runtime_reproducible.sh builds twice and requires the same archive. It did here: two independent builds gave identical bytes.

Qualification

It was checked on an NVIDIA RTX 3090, with Qwen3.5-9B Q4_K_M (SHA-256 cd76ec20…2a13), a 32,768-token context, flash attention and thinking off. The test set was 33 labeled invoices, run twice:

  • 33 of 33 documents completed in each run, at a mean of 13.1 and 13.4 seconds per document, with the slowest at 41.1 s;
  • 271 of 273 line items were exact in each run;
  • the two runs' answers were identical on 33 of 33;
  • the answers were identical, on 33 of 33, to those of the earlier build this one replaces;
  • peak graphics memory for the server was 6,393 MiB.

License

llama.cpp is MIT-licensed; its license is in LICENSE. The build recipe in recipe/ is also offered under the MIT license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support