Yi-Coder-9B-Chat-GGUF / Yi-Coder-9B-Chat.speed.md
roenb's picture
Upload folder using huggingface_hub
b6e81a2 verified
|
Raw
History Blame Contribute Delete
952 Bytes
# Speed benchmark — CPU, NVIDIA GeForce RTX 3090, NVIDIA RTX A4000, NVIDIA RTX A4000
_2026-07-18T12:30:55_ · averaged over the runs shown.
**Per-GPU comparison** lets you see newer vs older architecture throughput side by side (e.g. Ampere RTX 3090 vs Pascal GTX 1080 Ti). Higher = faster generation.
| File | Size | CPU gen t/s | NVIDIA GeForce RTX 3090 gen t/s | NVIDIA RTX A4000 gen t/s | NVIDIA RTX A4000 gen t/s |
| --- | --- | --- | --- | --- | --- |
| Yi-Coder-9B-Chat-Q4_K_M.gguf | 5082.1 MB | 7.8 | 120.1 | 64.2 | 64.8 |
| Yi-Coder-9B-Chat-Q5_K_M.gguf | 5968.3 MB | 6.8 | 107.9 | 56.4 | 56.9 |
| Yi-Coder-9B-Chat-Q6_K.gguf | 6910.0 MB | 5.9 | 93.2 | 47.1 | 48.6 |
| Yi-Coder-9B-Chat-Q8_0.gguf | 8949.2 MB | 4.8 | 80.5 | 40.2 | 40.3 |
Measured via llama-server on this machine (CPU = `-ngl 0`; each GPU pinned via
CUDA_VISIBLE_DEVICES with full offload). Generation = new-token throughput.
GPU numbers require a CUDA build of llama.cpp.