Image-to-Video
Diffusion Single File
LTX-2
text-to-video
video-to-video
image-text-to-video
audio-to-video
text-to-audio
video-to-audio
audio-to-audio
text-to-audio-video
image-to-audio-video
image-text-to-audio-video
ltx-video
lightricks
comfyui
ltx-2.5
Instructions to use Lightricks/LTX-2.5 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusion Single File
How to use Lightricks/LTX-2.5 with Diffusion Single File:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- LTX-2
How to use Lightricks/LTX-2.5 with LTX-2:
# Install the LTX-2 pipelines git clone https://github.com/Lightricks/LTX-2.git cd LTX-2 uv sync --extra natten
# Download weights from this repo # Substitute filenames from this repo's "Files and versions" if they differ hf download Lightricks/LTX-2.5 \ diffusion_models/<distilled-transformer>.safetensors \ text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ vae/<video-vae>.safetensors \ vae/<audio-vae>.safetensors \ latent_upscale_models/<spatial-upsampler>.safetensors \ latent_upscale_models/<temporal-upsampler>.safetensors \ --local-dir models/LTX-2.5 # DFR requires the detailing IC-LoRA (separate repo; strength is fixed at 0.5) hf download Lightricks/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler --local-dir models/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler# Distilled LTX-2.5 pipeline (fast) uv run python -m ltx_pipelines.distilled \ --transformer-path models/LTX-2.5/diffusion_models/<distilled-transformer>.safetensors \ --text-encoder-path models/LTX-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ --video-vae-path models/LTX-2.5/vae/<video-vae>.safetensors \ --audio-vae-path models/LTX-2.5/vae/<audio-vae>.safetensors \ --spatial-upsampler-path models/LTX-2.5/latent_upscale_models/<spatial-upsampler>.safetensors \ --num-frames 121 \ --prompt "A beautiful sunset over the ocean" \ --output-path output.mp4 # For image-to-video, add: --image path/to/image.jpg 0 0.8# DFR pipeline (higher detail fidelity; optional temporal 2x/4x) uv run python -m ltx_pipelines.dfr_pipeline \ --transformer-path models/LTX-2.5/diffusion_models/<distilled-transformer>.safetensors \ --text-encoder-path models/LTX-2.5/text_encoders/gemma4-12b-with-proj-ltx-2.5-bf16.safetensors \ --video-vae-path models/LTX-2.5/vae/<video-vae>.safetensors \ --audio-vae-path models/LTX-2.5/vae/<audio-vae>.safetensors \ --spatial-upsampler-path models/LTX-2.5/latent_upscale_models/<spatial-upsampler>.safetensors \ --temporal-upsampler-path models/LTX-2.5/latent_upscale_models/<temporal-upsampler>.safetensors \ --detailing-lora models/LTX-2.5-22b-IC-LoRA-Pixel-Spatial-Upscaler/ltx-2.5-22b-ic-lora-pixel-spatial-upscaler-x2-1.0.safetensors \ --spatial-upscalings 1 \ --temporal-upscalings 1 \ --height 1088 \ --width 1920 \ --num-frames 121 \ --prompt "A beautiful sunset over the ocean" \ --output-path output.mp4 # For 4K: --spatial-upscalings 2 --width 3840 --height 2176 # For image-to-video, add: --image path/to/image.jpg 0 0.8 - Notebooks
- Google Colab
- Kaggle
A2Vid two-stage (LTX-2.5): stage 2 distilled refine corrupts non-keyframe latents
#59
by bgyngell - opened
Setup
- Official A2VidPipelineTwoStage from LTX-2 (400fd31054597515f47125691032c04b1c3ee24e)
- Hub pack Lightricks/LTX-2.5: ltx-2.5-22b-dev-transformer-bf16 + Gemma TE + video/audio VAE + ltx-2.5-22b-distilled-lora-450-bf16 + spatial x2 upsampler
- Distilled LoRA on stage 2 only, LTXV_LORA_COMFY_RENAMING_MAP, strength 1.0
- 1536×1024, 24 fps, 30 steps, seed 10, detect_params video guider, DEFAULT_NEGATIVE_PROMPT
- Stereo speech WAV as A2Vid audio; original waveform muxed at the end (not VAE-decoded audio)
- Image conditioning: avatar still at frame 0, strength 1.0
- A100 80GB, offload_mode=cpu (LoRA fuse OOMs at ~80GB with offload_mode=none)
What we see
- Full two-stage: frame 0 is the pinned still. Later frames decode as a grid of coloured squares (not a face).
Latent dumps: - Stage 1 (1, 128, T, 16, 24): person visible on frame 0 and later frames
- After x2 upsample (1, 128, T, 32, 48): structure still looks like a person
- After stage 2: same spatial shape, but stats jump (~std 0.73 → 2.44, range ~±6 → ±14). Frame 0 still looks like a person; frame 1+ is salt-and-pepper. VAE decode of those frames is the colour mosaic.
- Mux/encode is not the cause: encoded MP4 stills match the bad decode.
- Skip stage 2 (decode the upsampled stage-1 latent with the same VAE): coherent talking-head clip, no lip motion. Read on this Stage 1 + upsample + decode look usable. Stage 2’s 3-step distilled Euler refine (STAGE_2_DISTILLED_SIGMAS + SimpleDenoiser + distilled LoRA) is what blows up every latent except the strength-1.0 keyframe.
The model files were downloaded directly from HF. I checked that the deployed files match the size of the files on HF.
Has anyone else had this issue? Can anyone help me figure out what is going wrong?