Commit History

Remove redundant w8_bf16 engines (int8 weights fold to bf16 at build → identical size/speed to the bf16 baseline)
63cc70d
verified

cortexelus commited on

Add sm_120 TRT engines (AOT SWA, wide profile) for the quantized tiers
4e8cc21
verified

cortexelus commited on

sm_120 SAME-L: AOT (graph-capturable) engines + tensorRT/README.md (#6)
a05f2d4

cortexelus commited on

Rebuild medium DiT fp16-mixed engines (sm_90 + new sm_120): attention now fuses, 4.3x faster (#4)
06debc3

cortexelus commited on