Sol-LTX-Infer demo videos (B300 / sm_103, CUDA 13)
Generated on a single NVIDIA B300 with the SGLang multimodal_gen stack. Steady-state (warmup-excluded) speedups:
| model | baseline | fullopt | fullopt + real NVFP4 |
|---|---|---|---|
| LTX-2.3 (1080p/10s, 2-stage HQ) | 99.3s | 44.3s (2.24x) | 41.4s (2.40x) |
| Cosmos3-Super 64B (720p/189f, 1 GPU) | 345.0s | 170.1s (2.03x) | 130.3s (2.65x) |
| SANA-Video 2B (480p) | 41.1s | (EasyCache+fusion+compile) | — |
fullopt = cache/sparse/fusion stack; +NVFP4 = non-TE shape-aware NVFP4 GEMM
(jit CUTLASS / autotuned flashinfer per shape). See repo EXECUTION_LOG.md.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support