# Speed benchmark — CPU, NVIDIA GeForce RTX 3090, NVIDIA RTX A4000, NVIDIA RTX A4000 _2026-07-18T12:30:55_ · averaged over the runs shown. **Per-GPU comparison** lets you see newer vs older architecture throughput side by side (e.g. Ampere RTX 3090 vs Pascal GTX 1080 Ti). Higher = faster generation. | File | Size | CPU gen t/s | NVIDIA GeForce RTX 3090 gen t/s | NVIDIA RTX A4000 gen t/s | NVIDIA RTX A4000 gen t/s | | --- | --- | --- | --- | --- | --- | | Yi-Coder-9B-Chat-Q4_K_M.gguf | 5082.1 MB | 7.8 | 120.1 | 64.2 | 64.8 | | Yi-Coder-9B-Chat-Q5_K_M.gguf | 5968.3 MB | 6.8 | 107.9 | 56.4 | 56.9 | | Yi-Coder-9B-Chat-Q6_K.gguf | 6910.0 MB | 5.9 | 93.2 | 47.1 | 48.6 | | Yi-Coder-9B-Chat-Q8_0.gguf | 8949.2 MB | 4.8 | 80.5 | 40.2 | 40.3 | Measured via llama-server on this machine (CPU = `-ngl 0`; each GPU pinned via CUDA_VISIBLE_DEVICES with full offload). Generation = new-token throughput. GPU numbers require a CUDA build of llama.cpp.