thread13 commited on
Commit
b042ad1
·
verified ·
1 Parent(s): 9c45e20
Files changed (1) hide show
  1. README.md +1 -1
README.md CHANGED
@@ -38,7 +38,7 @@ All uploaded models were tested and add a ~20-30 % gain in inference speed compa
38
 
39
  ### llama.cpp invocation examples - Kwaipilot_KAT-Coder-V2.5-Dev-IQ4_XS.bartowski.mtp.q8_k_xl.gguf
40
 
41
- * this fits 24 GiB VRAM via using quantized KV cache
42
 
43
 
44
  ```sh
 
38
 
39
  ### llama.cpp invocation examples - Kwaipilot_KAT-Coder-V2.5-Dev-IQ4_XS.bartowski.mtp.q8_k_xl.gguf
40
 
41
+ * this fits 24 GiB VRAM due to using quantized KV cache
42
 
43
 
44
  ```sh