FLUX.2 [klein] 4B for WebGPU

Quantized weights of FLUX.2 [klein] 4B, packed for text-to-image generation entirely in the browser with the Kaminos WebGPU inference kit. The page downloads about 3.2 GB once (gzip copies of everything except the embedding table, from which it fetches only the rows a prompt uses), caches it, and generates on the visitor's own GPU. No server-side compute is involved.

Try it in your browser · Source

In Chrome, a 512 × 512 image (4 steps) takes 6–7 seconds on an Apple M4 Max and about 20 seconds on a 16 GB Apple M2 Pro. The page also offers 768 and 1024, and frees each stage's working memory before the next, so a 16 GB Mac generates at 1024 × 1024 without swapping.

Contents

Folder Component Format Size
te/ Qwen3-4B text encoder, layers 0–26 (the pipeline reads hidden states 9, 18 and 27) int4 weights, group 128, affine; norms f16 1.45 GB
te/te-embed.bin Token-embedding table, read one row per token with HTTP range requests f16 0.78 GB
te/tokenizer.json Qwen2 byte-level BPE tokenizer JSON 11 MB
dit/ Rectified-flow transformer block linears int4 (group 128, affine); embedders, modulation and output projections int4 (group 64, affine); norms f16 2.07 GB
vae/ VAE decoder and latent batch-norm statistics f16; statistics f32 0.1 GB

Each folder has a manifest.json listing every tensor's shape, format, byte offset, and the SHA-256 of each file. Every bundle and the tokenizer also have a gzip copy (.gz, listed in the manifest), which is what the page downloads; the embedding table stays uncompressed for range requests. Quantization is weight-only. Activations and the residual stream stay in floating point.

Changes from the original

  • Weights quantized and repacked into per-block files as described above.
  • Text encoder truncated to the layers the FLUX.2 [klein] pipeline uses.
  • Projections that share an input are fused (for example query, key and value).
  • 3×3 VAE convolution weights reordered for channels-last execution.

License

Apache License 2.0, the same license as the original model. FLUX.2 [klein] 4B is by Black Forest Labs. Its text encoder is Qwen3-4B by the Qwen team, also Apache 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BasinShapers/flux2-klein-4b-webgpu

Quantized
(55)
this model