--- license: apache-2.0 tags: - ccsr - super-resolution - image-to-image - upscaling - controlnet - tensorrt - rtx - comfyui pipeline_tag: image-to-image --- # CCSR: TensorRT RTX Acceleration Engine Ultra-fast **TensorRT RTX** execution engine and auxiliary modules for **CCSR (Creative Content Super-Resolution)**, designed for real-time generative image upscaling in **ComfyUI**. > 📦 **ComfyUI Loader & Upscaler Extension:** > All nodes supporting TensorRT engine execution are available in: > 👉 **[https://github.com/ussoewwin/ComfyUI-NunchakuFluxLoraStacker](https://github.com/ussoewwin/ComfyUI-NunchakuFluxLoraStacker)** --- ## 🌟 Overview Creative Content Super-Resolution (CCSR) is a diffusion-based super-resolution framework leveraging a Controlled UNet and ControlNet structure to synthesize rich photorealistic textures and fine details. This repository provides an optimized **NVIDIA TensorRT RTX Engine** implementation for CCSR: - **Fused Denoising Engine (`ccsr_apply_f16io.rtxplan`)**: - Fuses the ControlNet and Controlled UNet denoising computation into a single compiled TensorRT engine. - Fixed 512px tile resolution (64×64 latent tile) executing at ~**24 ms/step** (~**4.7× speedup** over PyTorch FP16 at ~113 ms/step on modern RTX GPUs). - Synchronized stream execution on current PyTorch CUDA streams to eliminate race conditions and deadlocks. - **Engine-Only Deployment (`ccsr_trt_aux.safetensors`)**: - Contains only the essential companion modules: FP16 AutoencoderKL (VAE encoder/decoder) and condition encoder. - Automatically loaded alongside the engine, eliminating the need to download large full checkpoints (~3.2 GB saved). --- ## 📦 Available Files | Filename | Description | Architecture / Components | File Size | Recommended Location | License | | :--- | :--- | :--- | :--- | :--- | :--- | | `ccsr_apply_f16io.rtxplan` | TensorRT Fused Denoising Engine | ControlNet + UNet fused RTX Engine (Tile 512px / Latent 64×64) | ~1.4 GB | `custom_nodes/.../nodes/CCSR/trt_engines/` | Apache-2.0 | | `ccsr_trt_aux.safetensors` | TRT Auxiliary Weights | FP16 VAE AutoencoderKL + Condition Encoder | ~450 MB | `custom_nodes/.../nodes/CCSR/trt_engines/` | Apache-2.0 | --- ## ⚙️ Performance & Benchmark Comparison Measurements conducted on an NVIDIA RTX 4090 / RTX 5090 environment: | Execution Mode | Files Required | VRAM Overhead (Denoising) | Step Latency (Tile 512) | Speedup | | :--- | :--- | :--- | :--- | :--- | | **Stock CCSR (FP16 PyTorch)** | Full Checkpoint (~3.2 GB) | ~3.8 GiB | ~113 ms / step | 1.0× (Baseline) | | **CCSR TensorRT RTX** | Engine + Aux (~1.85 GB total) | ~2.2 GiB | ~**24 ms / step** | **~4.7× faster** | --- ## 🚀 Usage in ComfyUI TensorRT engine execution for CCSR is integrated natively into the **[ComfyUI-NunchakuFluxLoraStacker](https://github.com/ussoewwin/ComfyUI-NunchakuFluxLoraStacker)** custom-node pack. ### Workflow Example