Update README.md
Browse files
README.md
CHANGED
|
@@ -12,7 +12,7 @@ exllamav2 quantizations of TheDrummer's [Behemoth-R1-123B-v2](https://huggingfac
|
|
| 12 |
|
| 13 |
[2.50bpw h6](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/2.50bpw_H6) (Quantizing)
|
| 14 |
[4.25bpw h6](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/4.25bpw_H6) (61.324 GiB)
|
| 15 |
-
[8.00bpw h8](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/8.00bpw_H8) (114.559 GiB)
|
| 16 |
[measurement.json](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/resolve/main/measurement.json?download=true)
|
| 17 |
|
| 18 |
The 4.25bpw quant will squeeze into 3 24GB GPUs with 16k fp16 context, but can load with more than 64k context in 4 24GB GPUs.
|
|
|
|
| 12 |
|
| 13 |
[2.50bpw h6](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/2.50bpw_H6) (Quantizing)
|
| 14 |
[4.25bpw h6](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/4.25bpw_H6) (61.324 GiB)
|
| 15 |
+
[8.00bpw h8](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/8.00bpw_H8) (114.559 GiB) (Uploading)
|
| 16 |
[measurement.json](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/resolve/main/measurement.json?download=true)
|
| 17 |
|
| 18 |
The 4.25bpw quant will squeeze into 3 24GB GPUs with 16k fp16 context, but can load with more than 64k context in 4 24GB GPUs.
|