Instructions to use prithivMLmods/Qwen3.8-27B-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use prithivMLmods/Qwen3.8-27B-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="prithivMLmods/Qwen3.8-27B-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("prithivMLmods/Qwen3.8-27B-FP8") model = AutoModelForMultimodalLM.from_pretrained("prithivMLmods/Qwen3.8-27B-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use prithivMLmods/Qwen3.8-27B-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "prithivMLmods/Qwen3.8-27B-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Qwen3.8-27B-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/prithivMLmods/Qwen3.8-27B-FP8
- SGLang
How to use prithivMLmods/Qwen3.8-27B-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "prithivMLmods/Qwen3.8-27B-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Qwen3.8-27B-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "prithivMLmods/Qwen3.8-27B-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "prithivMLmods/Qwen3.8-27B-FP8", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use prithivMLmods/Qwen3.8-27B-FP8 with Docker Model Runner:
docker model run hf.co/prithivMLmods/Qwen3.8-27B-FP8
Qwen3.8-27B-FP8
Qwen3.8-27B-FP8 is a FP8 dynamic quantized version of Qwen/Qwen3.8-27B, a 27-billion-parameter dense causal language model with a native vision encoder from the Qwen team. Built on the Qwen3.5 architectural foundation, it is a compact, deployment-friendly member of the Qwen3.8 generation. This quantization reduces model size and memory footprint while preserving long-form reasoning, mathematical problem solving, scientific analysis, coding, multimodal understanding, and instruction-following capabilities, making deployment more accessible on smaller GPUs. Qwen3.8-27B features a 64-layer hybrid architecture that interleaves Gated DeltaNet linear-attention blocks with periodic Gated Attention layers, trained with Multi-Token Prediction. It supports a native 262,144-token context window (extensible to 1M via YaRN scaling), native image and video understanding (from STEM diagrams to hour-scale videos), and flexible thinking control via a
reasoning_effortparameter (xhigh/medium/low) with thinking enabled by default and historical reasoning preserved across turns. It delivers substantial gains over its predecessor Qwen3.6-27B and is competitive with larger models on agentic coding, computer/browser/mobile-use tasks, and multimodal tool use, while remaining strong on general reasoning benchmarks.
This model is an experimental release and may generate unexpected behaviors or reasoning artifacts in certain scenarios. Quantization to FP8 may introduce minor numerical differences relative to the bf16 source model.
Model Size Comparison
| Variant | Approximate Size |
|---|---|
| Qwen/Qwen3.8-27B (Full Precision BF16 / FP16) | ~54 GB |
| prithivMLmods/Qwen3.8-27B-FP8 (Compressed FP8) | ~36 GB |
Quantization Details
Quantization was performed using llmcompressor with the following recipe:
default_stage:
default_modifiers:
QuantizationModifier:
targets: [Linear]
ignore: ['re:.*lm_head', 're:.*embed_tokens$', 're:.*visual.*', 're:.*model.visual.*',
're:.*linear_attn.*']
scheme: FP8_DYNAMIC
bypass_divisibility_checks: false
requires_calibration_data: false
Linear layers are quantized to FP8 with dynamic per-tensor activation scaling, so no calibration dataset is required (requires_calibration_data: false). The lm_head, embedding table, any vision-tower (visual) components, and linear_attn layers are excluded from quantization and remain at full precision to preserve output-head fidelity and numerical stability.
| Base model | Qwen/Qwen3.8-27B |
| Quantization scheme | FP8_DYNAMIC (Linear layers only) |
| Format | compressed-tensors |
| Calibration data required | No (dynamic activation scaling) |
| Excluded from quantization | lm_head, embed_tokens, visual (if present), linear_attn |
Use with vLLM
Qwen3.8-27B-FP8 is served through vLLM with native support for compressed-tensors FP8 checkpoints.
Requirements
torch >= 2.11.0vllm >=0.27.1- A GPU with FP8 support recommended (Hopper or Blackwell class) for best throughput; also runs on Ampere with FP8 dequantized on the fly.
Serve
vllm serve prithivMLmods/Qwen3.8-27B-FP8 \
--max-model-len 32768
Client request
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
messages = [
{
"role": "user",
"content": "Explain how a transformer model processes text."
}
]
response = client.chat.completions.create(
model="prithivMLmods/Qwen3.8-27B-FP8",
messages=messages,
temperature=0.0,
max_tokens=512,
)
print(response.choices[0].message.content)
Quick Start with Transformers
pip install transformers accelerate
from transformers import AutoTokenizer, AutoModelForCausalLM
import torch
model = AutoModelForCausalLM.from_pretrained(
"prithivMLmods/Qwen3.8-27B-FP8",
torch_dtype="auto",
device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
"prithivMLmods/Qwen3.8-27B-FP8"
)
messages = [
{
"role": "user",
"content": "Explain how a transformer model processes text."
}
]
inputs = tokenizer.apply_chat_template(
messages,
tokenize=True,
add_generation_prompt=True,
return_tensors="pt"
).to(model.device)
outputs = model.generate(
inputs,
max_new_tokens=512
)
print(
tokenizer.decode(
outputs[0][inputs.shape[-1]:],
skip_special_tokens=True
)
)
Acknowledgments
- llmcompressor: Used to produce the FP8 dynamic quantization for this release.
- vLLM: Recommended inference engine with native compressed-tensors FP8 support.
- Downloads last month
- -
Model tree for prithivMLmods/Qwen3.8-27B-FP8
Base model
Qwen/Qwen3.8-27B