void-model-mlx-q8 / README.md
dgrauet's picture
Rename the build note key so it cannot collide with the per-component notes table
a960db4 verified
|
Raw
History Blame Contribute Delete
1.79 kB
metadata
library_name: mlx
license: apache-2.0
base_model: netflix/void-model
tags:
  - mlx
  - mlx-forge
  - apple-silicon
  - safetensors
  - quantized
  - int8

dgrauet/void-model-mlx-q8

Int8 quantization (group_size 64, transformer Linear weights only) of dgrauet/void-model-mlx, the MLX conversion of netflix/void-model.

Quantized with mlx-forge (mlx-forge convert void-model --quantize --bits 8).

Good quality/memory balance (~48 GB RAM recommended for the full two-pass pipeline). On 32 GB Macs use the q4 variant instead.

Usage

These weights can be used with void-model-mlx:

python -m void_mlx.infer \
    --sample sample/BigBen \
    --pass1 weights/q8/void_pass1.safetensors \
    --pass2 weights/q8/void_pass2.safetensors \
    --base-model /path/to/CogVideoX-Fun-V1.5-5b-InP-mlx-q8 \
    --steps 30 --max-frames 13 --height 352 --width 624 \
    --output result.gif

Keep quantize_config.json next to the weights (the loader also infers bits/group_size from the weight shapes if it is missing).

Related Projects

Files

  • config.json (365.00 B)
  • quantize_config.json (63.00 B)
  • split_model.json (1.23 KB)
  • void_pass1.safetensors (6.22 GB)
  • void_pass2.safetensors (6.22 GB)