v3 full-model card
46537e0 verified - mtp-draft Full standalone model: FP8 attention/shared + NVFP4 dense + GPTQ-MXFP4 experts + MTP draft + stitched index
- run-logs run logs
- vllm_overlay Full standalone model: FP8 attention/shared + NVFP4 dense + GPTQ-MXFP4 experts + MTP draft + stitched index
- 1.57 kB Full standalone model: FP8 attention/shared + NVFP4 dense + GPTQ-MXFP4 experts + MTP draft + stitched index
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 20.6 kB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration
- 5.13 GB GPTQ-calibrated MXFP4 routed experts, 75 layers, damp=1.0, 240K-token calibration