--- license: apache-2.0 tags: - quantization - sglang - cuda - inference - escha - qwen3 - dense library_name: sglang --- # Escha Runtime — `qwen3dense` By **[Escha Labs Inc.](https://eschalabs.com/)** The serving runtime for **Escha** 2-/3-bit (`escha`) quantized models of the **`qwen3_5` dense architecture** (Qwen3.8-27B and siblings). One repo per model architecture, one directory per engine — this architecture currently has **one** engine, [`sglang/`](sglang/). | | **SGLang** — [`sglang/`](sglang/) | |---|---| | Best for | everything: single user, teams, agents | | Concurrency | continuous batching, paged KV, optional radix prefix cache | | Tool calls / JSON schema / thinking parser | yes | | Interface | OpenAI-compatible (`/v1/chat/completions`, `/v1/completions`, `/v1/models`) | | Install | Python 3.12 venv + CUDA-12 PyTorch, then one wheel | The engine is a fork of [SGLang](https://github.com/sgl-project/sglang) bundled inside the wheel, running the Escha CUDA kernels. No separate `sglang` install is needed, and none should be present — the wheel ships its own. ### Compatible models | Model repo | Bits | |---|---| | [EschaLabs/Qwen3.8-27B-Escha-W2](https://huggingface.co/EschaLabs/Qwen3.8-27B-Escha-W2) | 2-bit, mixed-rate (`escha`) | > This runtime targets the **`qwen3_5` dense** architecture. Its wheel also happens to register the > `eschamoe` mixture-of-experts method, so a `qwen3_5_moe` model will load too — but the tuning, > the defaults in `sglang/serve.sh` and the documentation here are all written for the dense > architecture. For a mixture-of-experts model use > [`escha-runtime-qwen3moe`](https://huggingface.co/EschaLabs/escha-runtime-qwen3moe), whose > defaults are measured on it. A model of a genuinely different architecture will not load — use > the matching `escha-runtime-` repo. ## Quickstart Full detail, including the per-GPU cookbook and troubleshooting: [`sglang/INSTALL.md`](sglang/INSTALL.md). ```bash python3.12 -m venv .venv && source .venv/bin/activate pip install -U pip wheel pip install "torch==2.9.*" --index-url https://download.pytorch.org/whl/cu128 # cu12 torch FIRST pip install ./sglang/escha-*.whl # pulls the bundled sglang fork + its full dep closure hf download EschaLabs/Qwen3.8-27B-Escha-W2 --local-dir ./Qwen3.8-27B-Escha-W2 MODEL=./Qwen3.8-27B-Escha-W2 bash sglang/serve.sh ``` Then check the stack and the endpoint: ```bash python -c "import torch, escha, sglang; print(torch.cuda.is_available(), hasattr(torch.ops.escha, 'escham_decode_gemv'), escha.__version__)" curl -s http://127.0.0.1:30000/v1/models | python3 -m json.tool ``` > `pip install "torch==2.9.*"` is a hard pin, not a suggestion. A bare `torch>=2.9` resolves to a > newer minor and `import escha` then fails with `undefined symbol: _ZN3c10...` — the compiled > extension is ABI-linked to libtorch, and that ABI is not stable across PyTorch minors. ## Thinking, and why you probably want a budget This is a reasoning model. With thinking on, the reasoning arrives in `reasoning_content` and the answer in `content` — **read both**, or you will see half the response. Two per-request levers, both inside `chat_template_kwargs` (a *top-level* `enable_thinking` field is silently ignored): ```jsonc { "chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"} } ``` `reasoning_effort` is `"xhigh"` (**the default**), `"medium"` or `"low"`; anything else makes the template raise, which surfaces as an HTTP 400 rather than a silent fallback. It works by injecting one sentence of system instruction — `xhigh` asks the model to validate assumptions and weigh alternatives, `low` asks it to keep thinking brief, and **`medium` injects nothing at all**, so `medium` is the neutral, unsteered model rather than a midpoint. It therefore *asks* for shorter reasoning; it does not bound it. If you are running a benchmark or an agent, set a **thinking budget** instead, which forces `` after N reasoning tokens so an answer is always produced: see [`sglang/INSTALL.md` → Bounded thinking](sglang/INSTALL.md#bounded-thinking-thinking_budget) and [`sglang/thinking_budget.py`](sglang/thinking_budget.py). Without one, the usual failure is `finish_reason: "length"` with `content: null`, which a harness scores as *wrong* rather than as *slow*. ## Requirements - **NVIDIA GPU, compute capability 8.0–12.0** (Ampere → Blackwell), Linux x86-64 with **glibc ≥ 2.28**. The kernel launch route auto-selects per GPU; you never set it. Per-architecture and per-VRAM launch recipes: [`sglang/INSTALL.md` → Running on your GPU](sglang/INSTALL.md#running-on-your-gpu). - **Python 3.12** (the wheel is `cp312`-only) + **CUDA-12 PyTorch 2.9.x**. The wheel handles every other dependency. - **A working host C compiler and Python dev headers.** Triton JIT-compiles a small shim at CUDA-graph capture time — this is separate from `ptxas` and from a CUDA toolkit, so "driver only" does not cover it. On slim container images a stripped `libisl` breaks `cc1` while `gcc --version` still succeeds, and the failure surfaces ~40 s in as a `gcc` `CalledProcessError` inside `cuda_graph_runner.py` — which reads like a runtime bug and is not. Preflight in [`sglang/INSTALL.md`](sglang/INSTALL.md#requirements). - **24 GB VRAM** for the shipped defaults (65,536-token context, ~8–9 concurrent streams *at short prompts*) with a ~10.15 GB model. Those two figures are **not simultaneous**: the default pool is 68,686 tokens, which is one full-length 64k request or ~8 requests of ~8k. Note also the stream ceiling: `MAMBA_RATIO=0.3` sizes the recurrent-state pool, which clamps `max_running_requests` to 8–9 on a 24 GB card, so the `12`/`16` entries in the default `CUDA_GRAPH_BS` are **dropped and never captured**. To serve more streams raise `MAXREQ`/`MAXMAMBA` with `MEM` — the throughput recipe is in the [model card](https://huggingface.co/EschaLabs/Qwen3.8-27B-Escha-W2#verified-configurations); for long context (128k measured on a 24 GB card) see [By VRAM](sglang/INSTALL.md#by-vram). 16 GB should fit at a reduced context — the cookbook has a recipe, but we have not run it. ## Changelog **1.2.0** (2026-08-21) — **tensor parallelism (`--tp-size N`) now works.** The escha parameter class pins its own weight loader, which meant sglang's TP slicing never ran and every rank kept the whole checkpoint (rank 0 died with `weight must have shape (dim, width)`). It now slices per rank, including the fused-on-disk GDN `in_proj_qkv`, which is split into its three sub-projections first. > **Single-GPU users are unaffected.** Every new code path is gated on > `world_size > 1`; at `--tp-size 1` the loader is byte-for-byte what 1.1.1 did. Verified > as an identical shard layout and byte-identical greedy output. > > **TP > 1 is new and lightly tested — treat it as experimental.** It was contributed and > validated by [@ginerJuanUdesa](https://github.com/ginerJuanUdesa/escha-tp-fix-qwen3dense) > on 2× RTX 3090 (symmetric 6.02 GB/rank, coherent greedy output). **We have one GPU and > could not reproduce it**, and no numerical equivalence check against `--tp-size 1` has > been run yet. If you use it for evaluation, sanity-check a benchmark against the > single-GPU numbers first. Note that a multi-rank all-reduce reorders float accumulation, > so TP > 1 output is not expected to match TP = 1 bit-for-bit even when correct. > > On Ampere/Ada/Hopper you can add `DETERMINISTIC=1` to remove that reduction-order > variance if you want a stricter comparison. **1.1.1** (2026-08-21) — `process_weights_after_loading` now takes the rank's device instead of a hardcoded `cuda:0`. The hardcode put every 2-bit buffer on `cuda:0` while the input tensor sat on the server's actual device, so **any run not on device 0 — `--tp-size > 1`, or a single-GPU launch with `--base-gpu-id N` and no `CUDA_VISIBLE_DEVICES` — hit an illegal memory access on the first forward**, behind a traceback that pointed at the kernel rather than at the cause. Bit-identical wherever `cuda:0` was already correct, which is every configuration `serve.sh` ships. Reported with a diagnosis and a fix by [@ginerJuanUdesa](https://github.com/ginerJuanUdesa/escha-tp-fix-qwen3dense). **1.1.0** (2026-08-20) — first wheel with the dense (`escha`) serving path; the 1.0.x wheels registered `eschamoe` only, so a dense checkpoint failed at registry lookup. ## Known limitations - **Serving throughput has been measured on three cards** — RTX 5090 (32 GB, sm_120), RTX 4090 (24 GB, sm_89) and RTX 3090 (24 GB, sm_86); the 5090 and 3090 by independent evaluators working only from these docs. The 4090 is also where the model's thinking-on benchmarks (GPQA-Diamond, LiveCodeBench) were produced; its thinking-off commonsense suite ran on an L40. The **16 GB and 40 GB+** tiers in the cookbook remain configuration guidance derived from the model size and the wheel's architecture coverage, **not measurements**. Per-GPU numbers live on the [model card](https://huggingface.co/EschaLabs/Qwen3.8-27B-Escha-W2#performance-across-gpus). - **On Ampere the auto-selected kernel route is the slower one at batch 1.** `ESCHA_ROUTE` resolves to `lovelace` on sm_80/sm_86, but forcing `ESCHA_ROUTE=blackwell` measured **1.72× faster single-stream on an RTX 3090** (23.6 → 40.7 tok/s, TPOT 42.4 → 24.6 ms) with identical output. The two routes are bit-identical launch geometries, so this is safe to set; the gain is batch-1-only (parity at 2–16). Serving one user on Ampere? Set it. - **`DETERMINISTIC=1` fails on consumer Blackwell (sm_120).** The deterministic attention kernel requests 104 KB of shared memory per block, above the sm_120 limit, and the server exits during startup. It works on Ampere, Ada and Hopper. - **Greedy output is not bit-reproducible across requests.** Batch composition changes fp16 accumulation order, so a near-tie can flip and a long reasoning chain diverges from there. Two identical greedy requests may return different text. Use `DETERMINISTIC=1` when you need reproducibility, and never A/B two configurations by diffing one generation. - **CUDA-graph batch sizes are capped at 32**, the decode kernel's maximum M (`torch.ops.escha.escham_decode_gemv_max_m()`). The shipped default list stops at 16 because that is where aggregate throughput peaks on a 4090; capture at `24`/`32` works and is worth it if you serve that many streams. Past 32 a batch falls through to a large-M path meant for prefill, so the runtime refuses to capture it rather than bake in the wrong kernel. - **`ATTN_BACKEND=triton` is required on consumer Blackwell (RTX 50-series).** The default flashinfer backend asserts on this hybrid architecture at sm_120. The assertion names three acceptable backends — `triton`, `trtllm_mha`, `fa4` — of which only `triton` has been run on this model. Note that sm_120 shows steeper long-prompt decode decay than sm_89 (88.5% vs 96.3% of short-prompt rate at a 5,000-token prompt); the attention path is the obvious suspect and nobody has run the A/B that would confirm it. - **The model's stock chat template raises** on a few malformed conversations (a missing user query, a system message in the wrong position). That surfaces as an HTTP 400 from the template, not as a server fault. ## License Everything here is released under the **Apache License, Version 2.0** — see [`LICENSE`](LICENSE). All bundled third-party code is permissive (Apache-2.0 / MIT / BSD-3-Clause) — **no copyleft**. Full texts and the component inventory: [`THIRD_PARTY_LICENSES/`](THIRD_PARTY_LICENSES/). Model weights are **not** in this repo and carry their own license in the model repository.