SepGen generation checkpoint (gen-3k)

SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model

Aviad Dahan1, Rajaei Khatib1, Yonatan Bitton2, Idan Szpektor2, Lior Wolf1, Raja Giryes1

1Tel Aviv University, 2Google

Paper · Project page · Code · Other checkpoint: sep-12k (separation)

SepGen extends LTX-2.5 to emit the video, the audio-mix, and one waveform per captioned source in a single sampling run. The same weights generate (scene caption and per-source captions → video, audio-mix and stems) and separate (observed video and audio-mix → stems), selected by the noise level of the audio-mix.

gen-3k is the generation checkpoint: sep-12k continued for 3,000 steps in which 70% of samples present the audio-mix at a noise level no higher than the stems'. From a scene caption and one caption per source it generates the video, the audio-mix and one waveform per source; it also separates, at parity with sep-12k.

Files

File Content
lora_weights.safetensors the LoRA adapter (sha256 a72b3c662ce8d6b6e29013bbedb797d48ee1c90e8b7d219af4a27329eb58515b)
config.json adapter and training summary
train_config.yaml the full training config

The adapter targets the audio stream of the LTX-2.5 22B dev transformer: rank and alpha 128, 480 modules, 327M parameters, bf16, ComfyUI key format (diffusion_model.). It is trained with attention gates (block-diagonal caption routing, the protected audio-mix span, video reading audio from the audio-mix only, Cross-Stem Attention Guidance) that must also be installed at inference, so run it with the SepGen code, not as a plain LoRA.

Usage

git clone https://github.com/AviadDahan/SepGen && cd SepGen
bash setup.sh && source activate.sh && bash download_weights.sh
python generate.py examples/generation_prompts.json --out-dir outputs/gen      # uses gen-3k
python separate.py --batch-manifest examples/separation_manifest.json --checkpoint gen-3k --out-dir outputs/sep

--checkpoint gen-3k downloads this repo on first use.

License

These weights are a Derivative of LTX-2.x and are distributed under the LTX-2.x Community License Agreement, including its use-based restrictions (Attachment A) and the LTX Acceptable Use Policy.

Citation

@article{dahan2026sepgen,
  title   = {SepGen: Multi-Stem Audio-Video Separation and Generation in a Single Model},
  author  = {Dahan, Aviad and Khatib, Rajaei and Bitton, Yonatan and Szpektor, Idan and Wolf, Lior and Giryes, Raja},
  journal = {arXiv preprint arXiv:2610.11361},
  year    = {2026}
}
Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AviadDahan/SepGen-Generation

Adapter
(34)
this model

Collection including AviadDahan/SepGen-Generation

Paper for AviadDahan/SepGen-Generation