kylesayrs commited on
Commit
411e3c2
·
verified ·
1 Parent(s): f6c4264

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +10 -1
README.md CHANGED
@@ -25,6 +25,8 @@ vllm serve RedHatAI/GLM-5.2-FP8-NVFP4 --tensor_parallel_size 4 --kv_cache_dtype=
25
 
26
  This model was created using [LLM Compressor](https://github.com/vllm-project/llm-compressor). The example script can be found in `examples/quantizing_moe/glm5_example.py` [[Example] GLM5.2 Example](https://github.com/vllm-project/llm-compressor/pull/2869). Quantizing the model with data parallelism and 6xA100 takes about 3 hours.
27
 
 
 
28
  ```python
29
  import torch
30
  from compressed_tensors.offload import init_dist
@@ -135,4 +137,11 @@ model.save_pretrained(SAVE_DIR, save_compressed=True)
135
  tokenizer.save_pretrained(SAVE_DIR)
136
 
137
  torch.distributed.destroy_process_group()
138
- ```
 
 
 
 
 
 
 
 
25
 
26
  This model was created using [LLM Compressor](https://github.com/vllm-project/llm-compressor). The example script can be found in `examples/quantizing_moe/glm5_example.py` [[Example] GLM5.2 Example](https://github.com/vllm-project/llm-compressor/pull/2869). Quantizing the model with data parallelism and 6xA100 takes about 3 hours.
27
 
28
+ <details><summary>LLM Compressor Creation Script</summary>
29
+
30
  ```python
31
  import torch
32
  from compressed_tensors.offload import init_dist
 
137
  tokenizer.save_pretrained(SAVE_DIR)
138
 
139
  torch.distributed.destroy_process_group()
140
+ ```
141
+ </details>
142
+
143
+ ## Evaluation ##
144
+
145
+ | Benchmark | `zai-org/GLM-5.2` | `RedHatAI/GLM-5.2-NVFP4-FP8` |
146
+ | - | - | - |
147
+ | GPQA-Diamond | 91.2 | 89.1 |