Instructions to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8") model = AutoModelForMultimodalLM.from_pretrained("LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
- SGLang
How to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8 with Docker Model Runner:
docker model run hf.co/LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
Qwen3.8-27B-Humanlike-Chat 2.0 (FP8)
A 27B model that texts like a person and still does the work, in FP8 for 32 to 48 GB cards. No system prompt needed.
Results, examples and how it was made are on the main card: Qwen3.8-27B-Humanlike-Chat-GGUF.
"Base" here always means huihui-ai/Huihui-Qwen3.8-27B-abliterated, an abliterated Qwen3.8-27B. It is not the official Qwen release.
vLLM
Tested with vLLM 0.27.1 on an H100:
vllm serve LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8 --served-model-name humanlike-2.0 \
--max-model-len 16384 --language-model-only --linear-backend cutlass \
--enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3
On the H100, --linear-backend cutlass uses vLLM's precompiled block-FP8 kernel and the server is up in about 3 minutes. vLLM 0.27.1's default kernel on that card builds code at startup and needs nvcc in the container. On other GPUs, drop --linear-backend cutlass and vLLM picks a kernel for your card. Thinking is on by default. Send chat_template_kwargs: {"enable_thinking": false} per request for the fastest replies. Sampling: temperature 1.0, top_p 0.95, top_k 20.
Format
Same layout as Qwen's official Qwen/Qwen3.8-27B-FP8: FP8 e4m3 weights in 128 x 128 blocks with BF16 scales, dynamic activation scaling. 407 linear layers are FP8: MLP, attention, the linear-attention projections and the MTP head's layers. These stay BF16: embeddings, lm_head, norms, the linear-attention gate tensors (in_proj_a, in_proj_b, A_log, dt_bias, conv) and the vision tower. Made from the BF16 weights with no calibration data.
VRAM
The weights take 27.6 GiB in vLLM. A 48 GB card fits long contexts. A 32 GB card fits shorter contexts; lower --max-model-len and raise --gpu-memory-utilization to 0.95 if startup reports no memory for the cache. For 24 GB, use IQ4_XS or Q4_K_M from the main card.
Checks
Same check as every file on the main card: 5 fresh test chats against the reference (the base plus the 2.0 adapter at runtime in BF16), served on vLLM 0.27.1, one H100 NVL.
| File | KL vs reference | Top-1 agreement | Greedy turns identical | Check fails (greedy / sampled) | Decode tok/s |
|---|---|---|---|---|---|
| BF16 | 0.0032 | 97.2% | 12 of 15 | 0 / 0 | 46 |
| FP8 | 0.0079 | 91.7% | 9 of 15 | 0 / 0 | 72 |
Live requests on the same server: "yo you up" with thinking off came back yeah what's up with no reasoning. With thinking on, the train question came back 17:35 with the reasoning in its own field. A weather request produced get_weather with {"city": "Lisbon", "unit": "celsius"}.
Commissions
Commissions open, DM codebottle on Discord (https://discord.com/users/320486798859960322). I build custom finetunes like this one: characters, product voices, distillation into smaller models, and domain or use-case specific models.
License
Apache-2.0, inherited from the upstream Qwen and Huihui releases.
- Downloads last month
- 91
Model tree for LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8
Base model
Qwen/Qwen3.8-27B