Qwen3.8-27B-Humanlike-Chat 2.0 (FP8)

A 27B model that texts like a person and still does the work, in FP8 for 32 to 48 GB cards. No system prompt needed.

Results, examples and how it was made are on the main card: Qwen3.8-27B-Humanlike-Chat-GGUF.

"Base" here always means huihui-ai/Huihui-Qwen3.8-27B-abliterated, an abliterated Qwen3.8-27B. It is not the official Qwen release.

vLLM

Tested with vLLM 0.27.1 on an H100:

vllm serve LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8 --served-model-name humanlike-2.0 \
  --max-model-len 16384 --language-model-only --linear-backend cutlass \
  --enable-auto-tool-choice --tool-call-parser qwen3_coder --reasoning-parser qwen3

On the H100, --linear-backend cutlass uses vLLM's precompiled block-FP8 kernel and the server is up in about 3 minutes. vLLM 0.27.1's default kernel on that card builds code at startup and needs nvcc in the container. On other GPUs, drop --linear-backend cutlass and vLLM picks a kernel for your card. Thinking is on by default. Send chat_template_kwargs: {"enable_thinking": false} per request for the fastest replies. Sampling: temperature 1.0, top_p 0.95, top_k 20.

Format

Same layout as Qwen's official Qwen/Qwen3.8-27B-FP8: FP8 e4m3 weights in 128 x 128 blocks with BF16 scales, dynamic activation scaling. 407 linear layers are FP8: MLP, attention, the linear-attention projections and the MTP head's layers. These stay BF16: embeddings, lm_head, norms, the linear-attention gate tensors (in_proj_a, in_proj_b, A_log, dt_bias, conv) and the vision tower. Made from the BF16 weights with no calibration data.

VRAM

The weights take 27.6 GiB in vLLM. A 48 GB card fits long contexts. A 32 GB card fits shorter contexts; lower --max-model-len and raise --gpu-memory-utilization to 0.95 if startup reports no memory for the cache. For 24 GB, use IQ4_XS or Q4_K_M from the main card.

Checks

Same check as every file on the main card: 5 fresh test chats against the reference (the base plus the 2.0 adapter at runtime in BF16), served on vLLM 0.27.1, one H100 NVL.

File KL vs reference Top-1 agreement Greedy turns identical Check fails (greedy / sampled) Decode tok/s
BF16 0.0032 97.2% 12 of 15 0 / 0 46
FP8 0.0079 91.7% 9 of 15 0 / 0 72

Live requests on the same server: "yo you up" with thinking off came back yeah what's up with no reasoning. With thinking on, the train question came back 17:35 with the reasoning in its own field. A weather request produced get_weather with {"city": "Lisbon", "unit": "celsius"}.

Commissions

Commissions open, DM codebottle on Discord (https://discord.com/users/320486798859960322). I build custom finetunes like this one: characters, product voices, distillation into smaller models, and domain or use-case specific models.

License

Apache-2.0, inherited from the upstream Qwen and Huihui releases.

Downloads last month
91
Safetensors
Model size
28B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for LessThanThreeAI/Qwen3.8-27B-Humanlike-Chat-2.0-FP8