zkshared / nvfp4 /README.md
zk-Armor's picture
Upload nvfp4/README.md with huggingface_hub
844d2d5 verified
|
Raw
History Blame Contribute Delete
1.39 kB
# NVFP4 Kernel for Candle
## Overview
This repository contains an NVFP4 (NVIDIA 4-bit Floating Point) kernel implementation
for the [Candle](https://github.com/huggingface/candle) ML framework.
## NVFP4 Format
- **Element format**: E2M1 (4-bit float: 1 sign, 2 exponent, 1 mantissa)
- **Representable values**: ±0, ±0.5, ±1, ±1.5, ±2, ±3, ±4, ±6
- **Block size**: 16 elements
- **Block scale**: E4M3 FP8 (8 bits per block)
- **Memory layout**: 9 bytes per 16 elements (8 data + 1 scale)
- **Compression**: ~7.1x vs FP32
## Files
- `candle-kernels/src/nvfp4.cu` - CUDA kernels for quantize, dequantize, matmul
- `candle-core/src/quantized/mod.rs` - GgmlDType enum with NVFP4 variant
- `candle-core/src/quantized/cuda.rs` - CUDA backend integration
- `candle-core/src/quantized/k_quants.rs` - BlockNvfp4 Rust type and GgmlType impl
- `examples/nvfp4-inference.rs` - Rust inference example
- `test_nvfp4.py` - Python test suite (PyTorch-based validation)
## GPU
Tested on NVIDIA RTX PRO 6000 Blackwell Server Edition (SM 12.0, MIG 2g.48gb).
## Test Results
- E2M1 round-trip: max error 0.25 (expected for 4-bit precision)
- E4M3 scale round-trip: max relative error 7.1% (3 mantissa bits)
- Block quantization: relative error < 16% across distributions
- MatMul (1x256x512): relative error 10.1%
- MatMul (4x1024x2048): relative error 9.4%
- Memory savings: 7.11x compression vs FP32