How to use from the
Use from the
LTX.io library
# Install the LTX-2 pipelines
git clone https://github.com/Lightricks/LTX-2.git
cd LTX-2
uv sync --frozen
# Download the weights from this repo, plus the Gemma text encoder
hf download BingoG/SpatialAV2AV --local-dir models/SpatialAV2AV
hf download google/gemma-3-12b-it-qat-q4_0-unquantized --local-dir models/gemma-3-12b
# Fast pipeline (distilled model, no distilled LoRA needed)
uv run python -m ltx_pipelines.distilled \
    --distilled-checkpoint-path models/SpatialAV2AV/<distilled-checkpoint>.safetensors \
    --spatial-upsampler-path models/SpatialAV2AV/<spatial-upsampler>.safetensors \
    --gemma-root models/gemma-3-12b \
    --prompt "A beautiful sunset over the ocean" \
    --output-path output.mp4
# For image-to-video, add: --image path/to/image.jpg 0 0.8
# HQ pipeline (two-stage, higher quality)
uv run python -m ltx_pipelines.ti2vid_two_stages_hq \
    --checkpoint-path models/SpatialAV2AV/<checkpoint>.safetensors \
    --distilled-lora models/SpatialAV2AV/<distilled-lora>.safetensors 0.8 \
    --spatial-upsampler-path models/SpatialAV2AV/<spatial-upsampler>.safetensors \
    --gemma-root models/gemma-3-12b \
    --prompt "A beautiful sunset over the ocean" \
    --output-path output.mp4
# For image-to-video, add: --image path/to/image.jpg 0 0.8

SpatialAV2AV — Weights

Full fine-tuned transformer weights for spatial (binaural / stereo) audio-video editing on top of LTX-2.3 (22B). The model learns source (pre-edit sound field + video) + camera-trajectory instruction → edit (target sound field + video), producing real 2-channel stereo audio (L≠R) whose spatial image follows the camera motion.

Trained on the BingoG/LTX dataset (source→edit pairs, same clip rendered under different camera trajectories via Binamix HRTF).

Repository layout

Each checkpoint lives in its own step_XXXXX/ folder with an INFO.md recording the training run, step, loss, format, and inference usage.

Folder Step Notes
step_04000/ 4000 / 20000 Latest & recommended — best tested checkpoint (run 2026.06.19-17.33.11)

⚠️ Important: the file is torch-pickle, NOT safetensors

The weight file is named model_weights_step_04000.safetensors but it is actually a torch.save (zip/pickle) file, not a real safetensors container. This is a known quirk of the trainer (it uses accelerator.save, which writes torch format under a .safetensors name).

  • safetensors.torch.load_file(...)fails with HeaderTooLarge.
  • ✅ Load with torch.load(...) and inject into the base LTX-2.3 transformer via load_state_dict(..., strict=False) (verified missing=0 unexpected=0).

The project's SpatialAV2AV_inference.py --full_weights <file> path already does this correctly.

Usage (inference)

# base model + text encoder are the stock LTX-2.3 assets (not in this repo)
python packages/ltx-trainer/scripts/SpatialAV2AV_inference.py \
  --checkpoint  <LTX-2.3/ltx-2.3-22b-dev.safetensors> \
  --full_weights step_04000/model_weights_step_04000.safetensors \
  --text_encoder_path <gemma-3-12b-it-...> \
  --meta_json <...>+pan_right.json \
  --output out_pan_right.mp4 \
  --num_frames 113 --min_size 288 --steps 20 --guidance_scale 4.0 --seed 42
# prints "双声道 L≠R? True" on success

Download just one checkpoint folder:

from huggingface_hub import snapshot_download
d = snapshot_download("BingoG/SpatialAV2AV", allow_patterns="step_04000/*")
Downloads last month
26
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support