xocialize's picture
NOTICE: NVIDIA OML attribution + statement of quantization changes
c431224 verified
Raw
History Blame Contribute Delete
1.25 kB
Licensed by NVIDIA Corporation under the NVIDIA Open Model License
This repository redistributes a QUANTIZED derivative of the model weights of
nvidia/Qwen-Image-Flash (https://huggingface.co/nvidia/Qwen-Image-Flash). Use of these
weights is governed by the NVIDIA Open Model License Agreement:
https://www.nvidia.com/en-us/agreements/enterprise-software/nvidia-open-model-license
Additional Information: Apache License 2.0 (see LICENSE), as provided by the upstream
repository.
Qwen-Image-Flash is a DMD2 distillation of Qwen/Qwen-Image
(https://huggingface.co/Qwen/Qwen-Image).
Changes relative to upstream:
- transformer/model-int8.safetensors: the 20.43B DiT quantized to int8 (attention,
feed-forward and modulation linears, group size 64). The in/out projections, timestep
embedding and final norm are left at full precision.
- text_encoder/model-int8.safetensors: the Qwen2.5-VL-7B LANGUAGE MODEL quantized to
int8 (group size 64). The vision tower is omitted — text-to-image conditions on text
alone and never invokes it.
- vae/: unquantized, unchanged.
- tokenizer/tokenizer.json: converted from the upstream slow-tokenizer files
(vocab.json + merges.txt) and verified to produce identical token ids.