| SRC trellis_2_bf16.safetensors |
| compute/passthrough dtype: torch.bfloat16 |
|
|
| QUANTIZE 840 layers (int8+convrot, absmax): |
| x30 gs256 (3072, 1024) model.imgNshape.blocks.N.cross_attn.to_kv |
| x30 gs256 (1536, 1536) model.imgNshape.blocks.N.cross_attn.to_out |
| x30 gs256 (1536, 1536) model.imgNshape.blocks.N.cross_attn.to_q |
| x60 gs256 (1536, 8192) model.imgNshape.blocks.N.mlp.mlp.N |
| x30 gs256 (1536, 1536) model.imgNshape.blocks.N.self_attn.to_out |
| x30 gs256 (4608, 1536) model.imgNshape.blocks.N.self_attn.to_qkv |
| x30 gs256 (3072, 1024) model.imgNshape_N.blocks.N.cross_attn.to_kv |
| x30 gs256 (1536, 1536) model.imgNshape_N.blocks.N.cross_attn.to_out |
| x30 gs256 (1536, 1536) model.imgNshape_N.blocks.N.cross_attn.to_q |
| x60 gs256 (1536, 8192) model.imgNshape_N.blocks.N.mlp.mlp.N |
| x30 gs256 (1536, 1536) model.imgNshape_N.blocks.N.self_attn.to_out |
| x30 gs256 (4608, 1536) model.imgNshape_N.blocks.N.self_attn.to_qkv |
| x30 gs256 (3072, 1024) model.shapeNtxt.blocks.N.cross_attn.to_kv |
| x30 gs256 (1536, 1536) model.shapeNtxt.blocks.N.cross_attn.to_out |
| x30 gs256 (1536, 1536) model.shapeNtxt.blocks.N.cross_attn.to_q |
| x60 gs256 (1536, 8192) model.shapeNtxt.blocks.N.mlp.mlp.N |
| x30 gs256 (1536, 1536) model.shapeNtxt.blocks.N.self_attn.to_out |
| x30 gs256 (4608, 1536) model.shapeNtxt.blocks.N.self_attn.to_qkv |
| x30 gs256 (3072, 1024) model.structure_model.blocks.N.cross_attn.to_kv |
| x30 gs256 (1536, 1536) model.structure_model.blocks.N.cross_attn.to_out |
| x30 gs256 (1536, 1536) model.structure_model.blocks.N.cross_attn.to_q |
| x60 gs256 (1536, 8192) model.structure_model.blocks.N.mlp.mlp.N |
| x30 gs256 (1536, 1536) model.structure_model.blocks.N.self_attn.to_out |
| x30 gs256 (4608, 1536) model.structure_model.blocks.N.self_attn.to_qkv |
| groupsizes: {256: 840} quantized params: 5.10B (~5.1 GB int8) |
|
|
| LEAVE AS-IS (140 weights): |
| x120 not-2d |
| x19 not-in-indexed-block |
| x1 ineligible-K |
| 100/840 ... model.img2shape.blocks.21.cross_attn.to_out gs=256 relerr=0.99% cos=0.99995 |
| 200/840 ... model.img2shape.blocks.8.mlp.mlp.0 gs=256 relerr=0.81% cos=0.99997 |
| 300/840 ... model.img2shape_512.blocks.2.self_attn.to_out gs=256 relerr=0.85% cos=0.99996 |
| 400/840 ... model.img2shape_512.blocks.7.cross_attn.to_kv gs=256 relerr=0.79% cos=0.99997 |
| 500/840 ... model.shape2txt.blocks.19.cross_attn.to_q gs=256 relerr=0.81% cos=0.99997 |
| 600/840 ... model.shape2txt.blocks.5.mlp.mlp.2 gs=256 relerr=0.91% cos=0.99996 |
| 700/840 ... model.structure_model.blocks.17.self_attn.to_qkv gs=256 relerr=0.81% cos=0.99997 |
| 800/840 ... model.structure_model.blocks.4.cross_attn.to_out gs=256 relerr=0.86% cos=0.99996 |
| DONE: quantized 840 layers, 4240 tensors, 17.7s -> trellis_2_int8_convrot.safetensors |
|
|
| === quant error (relerr = ||dequant-source|| / ||source||) === |
| mean 0.851% min 0.765% max 1.392% layers 840 |
| per groupsize: gs256: mean 0.851% max 1.392% (x840) |
| worst 8 layers: |
| 1.392% cos 0.99990 gs256 model.img2shape.blocks.11.cross_attn.to_out |
| 1.386% cos 0.99990 gs256 model.img2shape_512.blocks.11.cross_attn.to_out |
| 1.192% cos 0.99993 gs256 model.shape2txt.blocks.14.cross_attn.to_out |
| 1.171% cos 0.99993 gs256 model.shape2txt.blocks.12.cross_attn.to_out |
| 1.140% cos 0.99994 gs256 model.shape2txt.blocks.11.cross_attn.to_out |
| 1.121% cos 0.99994 gs256 model.shape2txt.blocks.7.cross_attn.to_out |
| 1.121% cos 0.99994 gs256 model.shape2txt.blocks.22.cross_attn.to_out |
| 1.108% cos 0.99994 gs256 model.structure_model.blocks.1.cross_attn.to_out |