Size versus FP8 Quantization

#1
by SkyMind - opened

Total size of FP8 is 70.7GB versus 73GB for internlm/Intern-S2-Mobius.

Examining safetensors, there are some FP8_E4M3 values, but given the size, is everything that should be quantized being quantized?

Intern Large Models org

Thank for your comment, we have updated the model.

Messi-Hua changed discussion status to closed

Sign up or log in to comment