--- license: apache-2.0 base_model: Qwen/Qwen3-32B base_model_relation: quantized library_name: epure-runtime pipeline_tag: text-generation language: - en tags: - exeaon - epure - compressed - quantized --- # Exeaon1-Kese-32B Compressed with E-PURE. Runs with the free [`epure-runtime`](https://github.com/ExeaonLM/epure-runtime), and **stays compressed in memory** — the dense weight is never assembled. > **Requires `epure-runtime>=0.2.4`.** Earlier versions load every model as > float16. Qwen3-32B is trained in bfloat16 and its activations exceed the > float16 range, so on 0.2.3 this model overflows to infinity partway through > the second layer and produces garbage on prompts of ordinary length. This is > not a soft warning: the failure is silent on short prompts and total on real > ones. > > ``` > pip install "epure-runtime>=0.2.4" > ``` | | | |---|---| | Base model | [Qwen/Qwen3-32B](https://huggingface.co/Qwen/Qwen3-32B) | | Size on disk | **16.42 GB** | | Compression | **3.99×** | | Format | `.ebin` | ## Quality Measured against the base model on the same GPU, the same harness version and the same `limit`. Not copied from anyone's README. | benchmark | Qwen3-32B | Exeaon1-Kese-32B | delta | |---|---|---|---| | ARC-Challenge | 57.33 | 57.00 | -0.33 | | ARC-Easy | 82.67 | 82.67 | +0.00 | | HellaSwag | 69.33 | 69.67 | +0.34 | | PIQA | 82.33 | 83.00 | +0.67 | | **mean** | **72.91** | **73.09** | **+0.17** | **Retention: 100.2%** of base mean accuracy. The compressed model scores fractionally above the original. That is not a gain — it is measurement noise at `limit=300`, where one item is worth 0.33 points. What it does mean is that the loss from compression is smaller than this benchmark can resolve. Scored at `limit=300`, so roughly 2–3 points of standard error per task. Read the mean, not any single row. ## Usage ```python pip install "epure-runtime>=0.2.4" ``` ```python import epure model, tok = epure.load("model.ebin", device="cuda") print(tok.decode(model.generate(**tok("Hello", return_tensors="pt").to("cuda"), max_new_tokens=50)[0])) ``` ## Notes on speed On a large GPU this runs slower than dense fp16. A vendor tensor-core GEMM is heavily tuned and a large card has bandwidth to spare, so trading compute for memory loses there. The win is fitting in less memory — running where the dense model does not fit at all, on a smaller card, or alongside more of them. Qwen3-32B needs roughly 61 GB in bfloat16. This container needs 16.42 GB on disk and stays packed in memory. ## Licence The runtime is Apache-2.0 and free. The compression method that produced this container is proprietary to Zenux Plimver Technologies LTD.