File size: 1,315 Bytes
e33ec6b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 | ---
license: apache-2.0
base_model: Qwen/Qwen3-Coder-Next
tags:
- knot
- fat-station
- distributed-inference
- qwen3next
- moe
---
# qwen3-coder-next.knot
Sovereign `.knot` encoding of **Qwen3-Coder-Next** (`arch=qwen3next`) for the
Gnosis fat-station runtime. Converted from GGUF with `gguf-to-knot.py`
(K-quant block passthrough, arch-aware config).
## Why a knot
`.knot` is a range-addressable weight container: a station fetches only the
tensors for the layers it owns, over HTTP, instead of loading a whole file.
That is what lets an A3B MoE this size run **sharded across Cloudflare
Containers** (12 GiB each) rather than needing one large host.
## Architecture
Hybrid Mamba-2 gated-DeltaNet SSM with periodic full attention, and a
sparse MoE FFN on every layer. Roughly 3B parameters are active per token.
Long context is cheap because only the periodic attention layers carry a
KV cache.
## Serving
Pipeline-parallel across N stations, each owning a contiguous layer range:
```
FAT_STATION_KNOT=<this knot url>
FAT_STATION_SHARDED=1
FAT_STATION_ROLE=entry|mid|exit
FAT_STATION_LAYERS=0..6
```
MoE shards cleanly along layers: a layer owns its entire expert set, so
top-k routing never needs a tensor held by another station.
Published by AFFECTIVELY. https://huggingface.co/forkjoin-ai
|