MikeRoz
/

Behemoth-R1-123B-v2-exl2

Text Generation

Model card Files Files and versions

MikeRoz commited on Aug 22, 2025

Commit

76b7bff

·

verified ·

1 Parent(s): 91190b2

Create README.md

Files changed (1) hide show

README.md +19 -0

README.md ADDED Viewed

	@@ -0,0 +1,19 @@

+---
+inference: false
+base_model: TheDrummer/Behemoth-R1-123B-v2
+base_model_relation: quantized
+tags:
+  - exl2
+library_name: exllamav2
+pipeline_tag: text-generation
+---
+exllamav2 quantizations of TheDrummer's [Behemoth-R1-123B-v2](https://huggingface.co/TheDrummer/Behemoth-R1-123B-v2)
+[2.50bpw h6](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/2.50bpw_H6) (Quantizing)
+[4.25bpw h6](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/4.25bpw_H6) (61.324 GiB)
+[8.00bpw h8](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/tree/8.00bpw_H8) (114.559 GiB)
+[measurement.json](https://huggingface.co/MikeRoz/Behemoth-R1-123B-v2-exl2/resolve/main/measurement.json?download=true)
+The 4.25bpw quant will squeeze into 3 24GB GPUs with 16k fp16 context, but can load with more than 64k context in 4 24GB GPUs.
+The 8.00bpw quant requires 6 24 GB GPUs (or equivalent)