JuliaML's picture
Upload folder using huggingface_hub
e1213a2 verified
|
Raw
History Blame Contribute Delete
3.55 kB
---
license: other
license_name: openmdw1.1-license
license_link: https://openmdw.ai/license/1-1/
base_model: nvidia/Cosmos3-Super-Text2Image-4Step
pipeline_tag: text-to-image
library_name: diffusers
tags:
- cosmos3
- text-to-image
- quantization
- int4
- w4a16
- sdnq
- mlx
- apple-silicon
- cuda
---
# Cosmos3 Super Text-to-Image 4-Step — INT4 G64 BF16
Independent transformer-only quantization of
[`nvidia/Cosmos3-Super-Text2Image-4Step`](https://huggingface.co/nvidia/Cosmos3-Super-Text2Image-4Step).
This is not an official NVIDIA release.
Base revision: `0573a4b26b8e15d13d416e51f4680c8bc8b8c33d`.
Runtime, verification and benchmark code:
[`gtrg55/cosmos3-quant-mlx-cuda`](https://github.com/gtrg55/cosmos3-quant-mlx-cuda).
## Format
| Property | Value |
| --- | --- |
| Quantization | weight-only signed symmetric INT4 (W4A16) |
| Group size | 64 |
| Compute / activations | BF16 |
| BF16 transformer I/O | `proj_in`, `proj_out` |
| Quantized matrices | 898 |
| Transformer safetensors | 35.70 GiB |
| SVD / Hadamard / dynamic allocation | disabled |
| Base sampler | original distilled 4-step schedule, guidance 1.0 |
The repository contains the quantized transformer. Tokenizer, VAE and scheduler
are loaded from the original NVIDIA model.
## Validation status
- **Apple MLX / M4 Max:** validated end-to-end at 1024×1024.
- **Checkpoint packing:** exact sample round-trip verified.
- **CUDA:** loader and test command are supplied in the companion code repo;
publish a CUDA performance claim only after completing the included RunPod test.
Reference M4 Max result: 75.43 seconds generation, 38.97 GiB MLX peak memory.
Hardware, thermals and software versions affect timings.
![INT4 reference output](examples/sdnq-int4-g64-bf16io-mlx.png)
## Download and run on Apple Silicon
```bash
hf download nvidia/Cosmos3-Super-Text2Image-4Step \
--revision 0573a4b26b8e15d13d416e51f4680c8bc8b8c33d \
--local-dir models/Cosmos3-Super-Text2Image-4Step
hf download JuliaML/Cosmos3-Super-Text2Image-4Step-INT4-G64-BF16 \
--local-dir checkpoints/Cosmos3-INT4
python mlx/run_pipeline.py \
--model models/Cosmos3-Super-Text2Image-4Step \
--checkpoint checkpoints/Cosmos3-INT4/transformer \
--prompt prompts/winter-spruce-tall-1024.json \
--latents benchmarks/latents/cpu-seed1143-1024x1024-bf16.safetensors \
--output outputs/int4.png \
--height 1024 --width 1024 --seed 1143
```
Runtime code and pinned requirements are maintained in the
[companion GitHub repository](https://github.com/gtrg55/cosmos3-quant-mlx-cuda).
## Reproducibility assets
`examples/` contains the exact structured prompt, fixed BF16 latent, generation
configuration, output image and run report. Important hashes:
- prompt SHA-256: `3271fe4b38b1882cde4d3c87a4d4ae260c568cf4ab2d35c5d0f32bb12b1ba936`
- latent SHA-256: `6504206b2452a23680cac73d5763e3796260f7582f40ff59e6613cdb33f0aca6`
- reference PNG SHA-256: `23980b0c076f30915c71ec515de8d1e0d3b18138a6b797d9edb683703278c121`
## Limitations
- INT4 can change composition and fine detail relative to BF16 even with fixed
initial latents.
- This is weight-only quantization, not integer activation inference.
- The included sample is one controlled prompt, not a comprehensive quality
evaluation.
- Safety behavior is inherited from the base pipeline and the user's runtime
configuration.
## License and attribution
Distributed under OpenMDW-1.1. See `LICENSE` and `NOTICE`. The original NVIDIA
notices and license must remain with redistributed model materials.