HCHs's picture
Accelerate remaining FP8 linears with direct scaled-mm
4d29b4b verified
|
Raw
History Blame Contribute Delete
1.47 kB
# Attribution notice
RivetCoder-9B-A4B combines the following sources:
- Host model: `LiquidAI/LFM2.5-2.6B`, revision
`654f9463ce32b05d0429d76fe1f580b27d4c1ac0`, under the LFM Open License v1.0.
- Expert donor: `zai-org/GLM-5.3-Flash`, revision
`3f1971b7b5f7a528c9c4ef6212c8785298a8c24a`, under the MIT License.
- Router sequence teacher: `qwen/qwen3.8-27b`, accessed through a local LM Studio
endpoint. No Qwen weights are included.
- Architectural reference: `Akahsizrr/fuse-1-Lite`, revision
`430c959e47556ae53fed18a9d97f7cf30876e6ff`.
- Fast-serving kernel reference: `kernels-community/finegrained-fp8`, whose
Triton grouped-FP8 and activation-quantization kernels are Copyright 2026 The
Hugging Face Inc. team and licensed under Apache License 2.0. RivetCoder's
adaptation adds per-output scales, TorchAO expert-bank packing, deterministic
Top-4 reduction, direct FP8 Linear dispatch using the checkpoint's existing
qdata/scales, and automatic Windows MSVC environment discovery.
Modifications and new work include GLM expert selection, a fixed tied
identity-Hadamard bridge, expert/router weight folding, per-layer Top-4 routing,
token gating, bounded residual scaling, Qwen-supervised routing-control training,
the custom `fuse_glm` Transformers implementation, and the inference-only
grouped-FP8 serving runtime shipped in this repository.
See `provenance/` for the complete local build record with personal filesystem
paths removed.