Expert-pruned GLM-5.3 in INT4 W4A16 (compressed-tensors) for 4x H200 serving. Built on cyankiwi/GLM-5.3-AWQ-INT4. Massmax criterion.
0xSero
0xSero
AI & ML interests
Quantizing, benchmarking, training, and building.
Recent Activity
published an article about 18 hours ago
The DGX Spark Handbook liked a model 4 days ago
AikidoSec/altar-1 updated a model 8 days ago
0xSero/GLM-5.3-Flash-EXL3-SparkOrganizations
0xSero Releases — sorted by date
Inkling-Small EXL3 Quantization Suite
Public EXL3 trellis quantizations of thinkingmachines/Inkling-Small from 2.0 through 5.0 bpw.
Gemma — REAP
REAP-pruned Gemma-4 MoE.
Proven REAPs
Benchmarked REAP checkpoints with >=500 all-time downloads. GLM/Qwen/MiniMax/DeepSeek/Kimi/gemma.
MiniMax — REAP
REAP-pruned & quantized MiniMax-M2.1 / M2.7.
Trinity — REAP
REAP-pruned & quantized Trinity-Large-Thinking.
Datasets — observations & calibration
REAP layerwise observations, calibration sets, and training data.
GLM-5.3 REAP EXL3 Quantization Suite
Expert-pruned GLM-5.3 (753B) in EXL3. Massmax criterion, routed experts quantized, sensitive layers BF16. Ladder 500B-661B.
Local AI Registry
Hardware, models, launch recipes, prices, raw speed sweeps, and comparable local AI benchmarks.
Qwen — REAP
REAP-pruned & quantized Qwen3.5 / 3.6 / Coder variants.
GLM — REAP
REAP-pruned & quantized GLM-4.x / 5 / 5.1 (+ Flash fine-tunes).
DeepSeek — REAP
REAP-pruned & quantized DeepSeek-V4-Flash / V3.2.
Nemotron — REAP
REAP-pruned & quantized NVIDIA Nemotron-3-Super.
Other models
Hy3, Kimi-K2.5, INTELLECT-3, NousCoder.
GLM-5.3 REAP W4A16 (Hopper)
Expert-pruned GLM-5.3 in INT4 W4A16 (compressed-tensors) for 4x H200 serving. Built on cyankiwi/GLM-5.3-AWQ-INT4. Massmax criterion.
GLM-5.3 REAP EXL3 Quantization Suite
Expert-pruned GLM-5.3 (753B) in EXL3. Massmax criterion, routed experts quantized, sensitive layers BF16. Ladder 500B-661B.
0xSero Releases — sorted by date
Local AI Registry
Hardware, models, launch recipes, prices, raw speed sweeps, and comparable local AI benchmarks.
Inkling-Small EXL3 Quantization Suite
Public EXL3 trellis quantizations of thinkingmachines/Inkling-Small from 2.0 through 5.0 bpw.
Qwen — REAP
REAP-pruned & quantized Qwen3.5 / 3.6 / Coder variants.
Gemma — REAP
REAP-pruned Gemma-4 MoE.
GLM — REAP
REAP-pruned & quantized GLM-4.x / 5 / 5.1 (+ Flash fine-tunes).
Proven REAPs
Benchmarked REAP checkpoints with >=500 all-time downloads. GLM/Qwen/MiniMax/DeepSeek/Kimi/gemma.
DeepSeek — REAP
REAP-pruned & quantized DeepSeek-V4-Flash / V3.2.
MiniMax — REAP
REAP-pruned & quantized MiniMax-M2.1 / M2.7.
Nemotron — REAP
REAP-pruned & quantized NVIDIA Nemotron-3-Super.
Trinity — REAP
REAP-pruned & quantized Trinity-Large-Thinking.
Other models
Hy3, Kimi-K2.5, INTELLECT-3, NousCoder.
Datasets — observations & calibration
REAP layerwise observations, calibration sets, and training data.