buley's picture
add qwen3-coder-next.knot (qwen3next fat-station knot)
e33ec6b verified
|
Raw
History Blame Contribute Delete
1.32 kB
metadata
license: apache-2.0
base_model: Qwen/Qwen3-Coder-Next
tags:
  - knot
  - fat-station
  - distributed-inference
  - qwen3next
  - moe

qwen3-coder-next.knot

Sovereign .knot encoding of Qwen3-Coder-Next (arch=qwen3next) for the Gnosis fat-station runtime. Converted from GGUF with gguf-to-knot.py (K-quant block passthrough, arch-aware config).

Why a knot

.knot is a range-addressable weight container: a station fetches only the tensors for the layers it owns, over HTTP, instead of loading a whole file. That is what lets an A3B MoE this size run sharded across Cloudflare Containers (12 GiB each) rather than needing one large host.

Architecture

Hybrid Mamba-2 gated-DeltaNet SSM with periodic full attention, and a sparse MoE FFN on every layer. Roughly 3B parameters are active per token. Long context is cheap because only the periodic attention layers carry a KV cache.

Serving

Pipeline-parallel across N stations, each owning a contiguous layer range:

FAT_STATION_KNOT=<this knot url>
FAT_STATION_SHARDED=1
FAT_STATION_ROLE=entry|mid|exit
FAT_STATION_LAYERS=0..6

MoE shards cleanly along layers: a layer owns its entire expert set, so top-k routing never needs a tensor held by another station.

Published by AFFECTIVELY. https://huggingface.co/forkjoin-ai