NVFP4 text encoder for MiniMax-H3 video generation: the Qwen3-VL encoder quantized to run in ComfyUI on a single card.
Qtum Foundation
community
AI & ML interests
None defined yet.
Recent Activity
Sub-4-bit GGUF quants of DeepSeek-V4: Flash-0731 (seven PPL-tested tiers + DSpark draft) and Pro-0813 (1.57T, factory FP4 master).
Qwen3 4B to 32B in every format we ship: imatrix GGUF (seven tiers each) for llama.cpp, plus FP8, AWQ and GPTQ for vLLM.
NVFP4 text encoder for MiniMax-H3 video generation: the Qwen3-VL encoder quantized to run in ComfyUI on a single card.
Qwen3.8-27B (dense, vision) imatrix GGUF: 18 tiers plus q8_0/f16 mmproj and an MTP draft. More Qwen3.8 sizes as they land.
Sub-4-bit GGUF quants of DeepSeek-V4: Flash-0731 (seven PPL-tested tiers + DSpark draft) and Pro-0813 (1.57T, factory FP4 master).
Kimi K3 (2.8T params, 896 experts) GGUF quants, both verified running on one 8xH100 node. Includes a sub-512 GiB tier.
Qwen3 4B to 32B in every format we ship: imatrix GGUF (seven tiers each) for llama.cpp, plus FP8, AWQ and GPTQ for vLLM.