File size: 11,881 Bytes
bc55caa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
36bb7bf
 
 
 
 
 
 
 
 
 
bc55caa
 
ca22b2f
 
e05fa1a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
ca22b2f
 
 
 
 
 
 
 
 
 
 
 
bc55caa
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
---
license: apache-2.0
tags:
- quantization
- sglang
- cuda
- inference
- escha
- qwen3
- dense
library_name: sglang
---

# Escha Runtime β€” `qwen3dense`

By **[Escha Labs Inc.](https://eschalabs.com/)**

The serving runtime for **Escha** 2-/3-bit (`escha`) quantized models of the **`qwen3_5` dense
architecture** (Qwen3.8-27B and siblings). One repo per model architecture, one directory per
engine β€” this architecture currently has **one** engine, [`sglang/`](sglang/).

| | **SGLang** β€” [`sglang/`](sglang/) |
|---|---|
| Best for | everything: single user, teams, agents |
| Concurrency | continuous batching, paged KV, optional radix prefix cache |
| Tool calls / JSON schema / thinking parser | yes |
| Interface | OpenAI-compatible (`/v1/chat/completions`, `/v1/completions`, `/v1/models`) |
| Install | Python 3.12 venv + CUDA-12 PyTorch, then one wheel |

The engine is a fork of [SGLang](https://github.com/sgl-project/sglang) bundled inside the wheel,
running the Escha CUDA kernels. No separate `sglang` install is needed, and none should be
present β€” the wheel ships its own.

### Compatible models

| Model repo | Bits |
|---|---|
| [EschaLabs/Qwen3.8-27B-Escha-W2](https://huggingface.co/EschaLabs/Qwen3.8-27B-Escha-W2) | 2-bit, mixed-rate (`escha`) |
<!-- add each qwen3_5 dense Escha model you publish here -->

> This runtime targets the **`qwen3_5` dense** architecture. Its wheel also happens to register the
> `eschamoe` mixture-of-experts method, so a `qwen3_5_moe` model will load too β€” but the tuning,
> the defaults in `sglang/serve.sh` and the documentation here are all written for the dense
> architecture. For a mixture-of-experts model use
> [`escha-runtime-qwen3moe`](https://huggingface.co/EschaLabs/escha-runtime-qwen3moe), whose
> defaults are measured on it. A model of a genuinely different architecture will not load β€” use
> the matching `escha-runtime-<arch>` repo.

## Quickstart

Full detail, including the per-GPU cookbook and troubleshooting:
[`sglang/INSTALL.md`](sglang/INSTALL.md).

```bash
python3.12 -m venv .venv && source .venv/bin/activate
pip install -U pip wheel
pip install "torch==2.9.*" --index-url https://download.pytorch.org/whl/cu128   # cu12 torch FIRST
pip install ./sglang/escha-*.whl   # pulls the bundled sglang fork + its full dep closure

hf download EschaLabs/Qwen3.8-27B-Escha-W2 --local-dir ./Qwen3.8-27B-Escha-W2
MODEL=./Qwen3.8-27B-Escha-W2 bash sglang/serve.sh
```

Then check the stack and the endpoint:

```bash
python -c "import torch, escha, sglang; print(torch.cuda.is_available(), hasattr(torch.ops.escha, 'escham_decode_gemv'), escha.__version__)"
curl -s http://127.0.0.1:30000/v1/models | python3 -m json.tool
```

> `pip install "torch==2.9.*"` is a hard pin, not a suggestion. A bare `torch>=2.9` resolves to a
> newer minor and `import escha` then fails with `undefined symbol: _ZN3c10...` β€” the compiled
> extension is ABI-linked to libtorch, and that ABI is not stable across PyTorch minors.

## Thinking, and why you probably want a budget

This is a reasoning model. With thinking on, the reasoning arrives in `reasoning_content` and the
answer in `content` β€” **read both**, or you will see half the response.

Two per-request levers, both inside `chat_template_kwargs` (a *top-level* `enable_thinking` field
is silently ignored):

```jsonc
{ "chat_template_kwargs": {"enable_thinking": true, "reasoning_effort": "xhigh"} }
```

`reasoning_effort` is `"xhigh"` (**the default**), `"medium"` or `"low"`; anything else makes the
template raise, which surfaces as an HTTP 400 rather than a silent fallback. It works by injecting
one sentence of system instruction β€” `xhigh` asks the model to validate assumptions and weigh
alternatives, `low` asks it to keep thinking brief, and **`medium` injects nothing at all**, so
`medium` is the neutral, unsteered model rather than a midpoint. It therefore *asks* for shorter
reasoning; it does not bound it. If you are running a benchmark
or an agent, set a **thinking budget** instead, which forces `</think>` after N reasoning tokens so
an answer is always produced: see
[`sglang/INSTALL.md` β†’ Bounded thinking](sglang/INSTALL.md#bounded-thinking-thinking_budget) and
[`sglang/thinking_budget.py`](sglang/thinking_budget.py). Without one, the usual failure is
`finish_reason: "length"` with `content: null`, which a harness scores as *wrong* rather than as
*slow*.

## Requirements

- **NVIDIA GPU, compute capability 8.0–12.0** (Ampere β†’ Blackwell), Linux x86-64 with
  **glibc β‰₯ 2.28**. The kernel launch route auto-selects per GPU; you never set it. Per-architecture
  and per-VRAM launch recipes:
  [`sglang/INSTALL.md` β†’ Running on your GPU](sglang/INSTALL.md#running-on-your-gpu).
- **Python 3.12** (the wheel is `cp312`-only) + **CUDA-12 PyTorch 2.9.x**. The wheel handles every
  other dependency.
- **A working host C compiler and Python dev headers.** Triton JIT-compiles a small shim at
  CUDA-graph capture time β€” this is separate from `ptxas` and from a CUDA toolkit, so "driver only"
  does not cover it. On slim container images a stripped `libisl` breaks `cc1` while
  `gcc --version` still succeeds, and the failure surfaces ~40 s in as a `gcc`
  `CalledProcessError` inside `cuda_graph_runner.py` β€” which reads like a runtime bug and is not.
  Preflight in [`sglang/INSTALL.md`](sglang/INSTALL.md#requirements).
- **24 GB VRAM** for the shipped defaults (65,536-token context, ~8–9 concurrent streams *at short
  prompts*) with a ~10.15 GB model. Those two figures are **not simultaneous**: the default pool is
  68,686 tokens, which is one full-length 64k request or ~8 requests of ~8k. Note also the stream
  ceiling: `MAMBA_RATIO=0.3` sizes the recurrent-state pool, which clamps `max_running_requests` to
  8–9 on a 24 GB card, so the `12`/`16` entries in the default `CUDA_GRAPH_BS` are **dropped and
  never captured**. To serve more streams raise `MAXREQ`/`MAXMAMBA` with `MEM` β€” the throughput
  recipe is in the
  [model card](https://huggingface.co/EschaLabs/Qwen3.8-27B-Escha-W2#verified-configurations); for
  long context (128k measured on a 24 GB card) see
  [By VRAM](sglang/INSTALL.md#by-vram).
  16 GB should fit at a reduced context β€” the cookbook has a recipe, but we have not run it.

## Changelog

**1.2.0** (2026-08-21) β€” **tensor parallelism (`--tp-size N`) now works.** The escha
parameter class pins its own weight loader, which meant sglang's TP slicing never ran and
every rank kept the whole checkpoint (rank 0 died with `weight must have shape
(dim, width)`). It now slices per rank, including the fused-on-disk GDN `in_proj_qkv`,
which is split into its three sub-projections first.

> **Single-GPU users are unaffected.** Every new code path is gated on
> `world_size > 1`; at `--tp-size 1` the loader is byte-for-byte what 1.1.1 did. Verified
> as an identical shard layout and byte-identical greedy output.
>
> **TP > 1 is new and lightly tested β€” treat it as experimental.** It was contributed and
> validated by [@ginerJuanUdesa](https://github.com/ginerJuanUdesa/escha-tp-fix-qwen3dense)
> on 2Γ— RTX 3090 (symmetric 6.02 GB/rank, coherent greedy output). **We have one GPU and
> could not reproduce it**, and no numerical equivalence check against `--tp-size 1` has
> been run yet. If you use it for evaluation, sanity-check a benchmark against the
> single-GPU numbers first. Note that a multi-rank all-reduce reorders float accumulation,
> so TP > 1 output is not expected to match TP = 1 bit-for-bit even when correct.
>
> On Ampere/Ada/Hopper you can add `DETERMINISTIC=1` to remove that reduction-order
> variance if you want a stricter comparison.

**1.1.1** (2026-08-21) β€” `process_weights_after_loading` now takes the rank's device
instead of a hardcoded `cuda:0`. The hardcode put every 2-bit buffer on `cuda:0` while
the input tensor sat on the server's actual device, so **any run not on device 0 β€”
`--tp-size > 1`, or a single-GPU launch with `--base-gpu-id N` and no
`CUDA_VISIBLE_DEVICES` β€” hit an illegal memory access on the first forward**, behind a
traceback that pointed at the kernel rather than at the cause. Bit-identical wherever
`cuda:0` was already correct, which is every configuration `serve.sh` ships. Reported
with a diagnosis and a fix by [@ginerJuanUdesa](https://github.com/ginerJuanUdesa/escha-tp-fix-qwen3dense).

**1.1.0** (2026-08-20) β€” first wheel with the dense (`escha`) serving path; the 1.0.x
wheels registered `eschamoe` only, so a dense checkpoint failed at registry lookup.

## Known limitations

- **Serving throughput has been measured on three cards** β€” RTX 5090 (32 GB, sm_120), RTX 4090
  (24 GB, sm_89) and RTX 3090 (24 GB, sm_86); the 5090 and 3090 by independent evaluators working
  only from these docs. The 4090 is also where the model's thinking-on benchmarks (GPQA-Diamond,
  LiveCodeBench) were produced; its thinking-off commonsense suite ran on an L40. The **16 GB and
  40 GB+** tiers in the cookbook remain configuration guidance derived from the model size and the
  wheel's architecture coverage, **not measurements**. Per-GPU numbers live on the
  [model card](https://huggingface.co/EschaLabs/Qwen3.8-27B-Escha-W2#performance-across-gpus).
- **On Ampere the auto-selected kernel route is the slower one at batch 1.** `ESCHA_ROUTE`
  resolves to `lovelace` on sm_80/sm_86, but forcing `ESCHA_ROUTE=blackwell` measured **1.72Γ—
  faster single-stream on an RTX 3090** (23.6 β†’ 40.7 tok/s, TPOT 42.4 β†’ 24.6 ms) with identical
  output. The two routes are bit-identical launch geometries, so this is safe to set; the gain is
  batch-1-only (parity at 2–16). Serving one user on Ampere? Set it.
- **`DETERMINISTIC=1` fails on consumer Blackwell (sm_120).** The deterministic attention kernel
  requests 104 KB of shared memory per block, above the sm_120 limit, and the server exits during
  startup. It works on Ampere, Ada and Hopper.
- **Greedy output is not bit-reproducible across requests.** Batch composition changes fp16
  accumulation order, so a near-tie can flip and a long reasoning chain diverges from there. Two
  identical greedy requests may return different text. Use `DETERMINISTIC=1` when you need
  reproducibility, and never A/B two configurations by diffing one generation.
- **CUDA-graph batch sizes are capped at 32**, the decode kernel's maximum M
  (`torch.ops.escha.escham_decode_gemv_max_m()`). The shipped default list stops at 16 because
  that is where aggregate throughput peaks on a 4090; capture at `24`/`32` works and is worth it if
  you serve that many streams. Past 32 a batch falls through to a large-M path meant for prefill,
  so the runtime refuses to capture it rather than bake in the wrong kernel.
- **`ATTN_BACKEND=triton` is required on consumer Blackwell (RTX 50-series).** The default
  flashinfer backend asserts on this hybrid architecture at sm_120. The assertion names three
  acceptable backends β€” `triton`, `trtllm_mha`, `fa4` β€” of which only `triton` has been run on
  this model. Note that sm_120 shows steeper long-prompt decode decay than sm_89 (88.5% vs 96.3%
  of short-prompt rate at a 5,000-token prompt); the attention path is the obvious suspect and
  nobody has run the A/B that would confirm it.
- **The model's stock chat template raises** on a few malformed conversations (a missing user
  query, a system message in the wrong position). That surfaces as an HTTP 400 from the template,
  not as a server fault.

## License

Everything here is released under the **Apache License, Version 2.0** β€” see [`LICENSE`](LICENSE).
All bundled third-party code is permissive (Apache-2.0 / MIT / BSD-3-Clause) β€” **no copyleft**.
Full texts and the component inventory:
[`THIRD_PARTY_LICENSES/`](THIRD_PARTY_LICENSES/). Model weights are **not** in this repo and carry
their own license in the model repository.