Instructions to use ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ") model = AutoModelForMultimodalLM.from_pretrained("ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ
- SGLang
How to use ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ with Docker Model Runner:
docker model run hf.co/ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ
Qwen3.8-27B — 3-bit GSQ
Low-bit quantization of Qwen/Qwen3.8-27B
produced with GSQ (Gumbel-Softmax Quantization).
This checkpoint applies GSQ-based post-training quantization to the model weights, reducing precision while preserving the original model's reasoning, coding, multilingual, long-context, and agentic capabilities.
The transformer weights are quantized to 3-bit GSQ with group size 128. The embedding layer and LM head are quantized separately to 4-bit RTN with group size 64 to preserve output quality and embedding fidelity.
- Paper: GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling (arXiv:2604.18556)
- Paper page on HF: https://huggingface.co/papers/2604.18556
- Code: https://github.com/IST-DASLab/GSQ
- Collection: https://huggingface.co/collections/ISTA-DASLab/gsq
Evaluation Results
We evaluate the quantized checkpoint against the original
Qwen/Qwen3.8-27B. Both models were evaluated with xhigh thinking enabled.
Accuracy
| Benchmark | Base Model | 3-bit GSQ |
|---|---|---|
| AIME 2025 | 100.00 | 100.00 |
| GPQA Diamond | 89.90 | 91.41 |
Generation Statistics
| Benchmark | Base Model | 3-bit GSQ |
|---|---|---|
| AIME 2025 | 0.603M | 0.615M |
| GPQA Diamond | 3.721M | 3.705M |
Note: Due to the stochastic nature of these tasks, benchmark results can exhibit significant variance across runs. The reported scores correspond to a single evaluation run and should not be interpreted as definitive estimates of model performance.
Calibration dataset
The GSQ quantization calibration dataset was constructed to represent a broad range of LLM workloads, including reasoning, coding, scientific tasks, multilingual understanding, long-context processing, and agentic behaviour.
The calibration mixture consists of:
| Category | Percentage |
|---|---|
| Math | 13.5% |
| Code | 17.5% |
| Science | 20.0% |
| General | 12.5% |
| Multilingual | 12.5% |
| Long context | 14.0% |
| Agentic trajectories | 10.0% |
Serving with vLLM
Serving this checkpoint requires a patched vLLM installation.
Requirements:
- vLLM 0.27.1
- The included patch:
patch_vllm_qwen35_embedding.py
This patch enables vLLM support for quantized embedding weights.
Install vLLM:
pip install vllm==0.27.1
Apply the patch in the same Python environment:
python patch_vllm_qwen35_embedding.py
Then serve the model:
vllm serve ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
Important: If vLLM is reinstalled or the environment is recreated, run the patch again before serving the checkpoint.
Model size / text-only usage
The full checkpoint size is approximately 11.83 GB when deployed with vision capabilities enabled.
The quantization calibration dataset used for this release did not include vision samples. Therefore, while the vision components are preserved in the checkpoint and can be loaded, they were not calibrated using multimodal calibration data.
For text-only deployment, the vision components are not required. The model can
be loaded with the --language-model-only option:
vllm serve ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ \
--enable-auto-tool-choice \
--tool-call-parser qwen3_coder \
--reasoning-parser qwen3 \
--language-model-only
This removes the vision-related components from the loaded model and reduces the checkpoint size to approximately 10.90 GB.
This release does not currently support speculative decoding. The MTP (Multi-Token Prediction) components have been removed from the published checkpoint and are not available for MTP-based inference.
Actual VRAM usage during serving will be higher than the raw checkpoint size and depends on:
- KV-cache allocation
- context length
- batch size
- vLLM runtime overhead
- GPU architecture
Citation
@article{gsq2026,
title = {GSQ: Highly-Accurate Low-Precision Scalar Quantization for LLMs via Gumbel-Softmax Sampling},
author = {Dadgarnia, Alireza and Tabesh, Soroush and Nikdan, Mahdi and Helcig, Michael and Kurti{\'c}, Eldar and Kleinegger, Max and Alistarh, Dan},
journal= {arXiv preprint arXiv:2604.18556},
year = {2026},
url = {https://arxiv.org/abs/2604.18556}
}
- Downloads last month
- 150
Model tree for ISTA-DASLab/Qwen3.8-27B-3Bit-GSQ
Base model
Qwen/Qwen3.8-27B