Instructions to use LiconStudio/LTX-2.5-Multiple-Subject-Reference with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Diffusers
How to use LiconStudio/LTX-2.5-Multiple-Subject-Reference with Diffusers:
pip install -U diffusers transformers accelerate
import torch from diffusers import DiffusionPipeline # switch to "mps" for apple devices pipe = DiffusionPipeline.from_pretrained("LiconStudio/LTX-2.5-Multiple-Subject-Reference", dtype=torch.bfloat16, device_map="cuda") prompt = "Astronaut in a jungle, cold color palette, muted colors, detailed, 8k" image = pipe(prompt).images[0] - Notebooks
- Google Colab
- Kaggle
identities not kept
I used the workflow: LTX2.5-MSR-sample-workflow-V2.json
and the following prompt:
Image 1: A man in jedi clothes, jedi robe and boots holding an orange lightsaber in one hand.
Image 2: A woman in jedi clothes, jedi robe and boots holding a green lightsaber in one hand.
Image 5: Scene, Tatooine planet from Star Wars. Binary sunset view with 2 suns.The camera axis remains fixed throughout the sequence. Image 1 is established on the left side of the frame and Image 2 on the right side. Preserve their faces, bodies, identities, clothing, screen positions across every hard cut.
full body shot of two person on the Tatooine planet matching Image 5. Image 1 is on the left side with his orange lightsaber. Image 2 is on the right side with her green lightsaber.
They approach eachother. Then the woman (image 2) laughs and says, "Okay, let's start" The man (image 1) grins, and says, "Alright, if that's what you want". They both raise and clash their lightsabers in the air and the sound of clashing lightsabers fill the area.
But the result is not even close:
- 2 characters are totally different from the characters in image 1 and image 2 (secene is correct from image 5)
- characters are not doing the correct act
- characters are talking in a language that I don't understand (Chinese ?)
- lightsabers are weird
Everything in the workflow is same as the original workflow from here.
I just changed the diffusion model weight from default to fp8_e5m2 because although I have 128GB unified ram, workflow stucks at "Requested to load LTXAV" if I use "default" diffusion model weight.
So what is wrong ?
I used the workflow: LTX2.5-MSR-sample-workflow-V2.json
and the following prompt:Image 1: A man in jedi clothes, jedi robe and boots holding an orange lightsaber in one hand.
Image 2: A woman in jedi clothes, jedi robe and boots holding a green lightsaber in one hand.
Image 5: Scene, Tatooine planet from Star Wars. Binary sunset view with 2 suns.The camera axis remains fixed throughout the sequence. Image 1 is established on the left side of the frame and Image 2 on the right side. Preserve their faces, bodies, identities, clothing, screen positions across every hard cut.
full body shot of two person on the Tatooine planet matching Image 5. Image 1 is on the left side with his orange lightsaber. Image 2 is on the right side with her green lightsaber.
They approach eachother. Then the woman (image 2) laughs and says, "Okay, let's start" The man (image 1) grins, and says, "Alright, if that's what you want". They both raise and clash their lightsabers in the air and the sound of clashing lightsabers fill the area.But the result is not even close:
- 2 characters are totally different from the characters in image 1 and image 2 (secene is correct from image 5)
- characters are not doing the correct act
- characters are talking in a language that I don't understand (Chinese ?)
- lightsabers are weird
Everything in the workflow is same as the original workflow from here.
I just changed the diffusion model weight from default to fp8_e5m2 because although I have 128GB unified ram, workflow stucks at "Requested to load LTXAV" if I use "default" diffusion model weight.So what is wrong ?
Hi,
Thanks for sharing your prompt, results, and hardware details.
We tested the revised action prompt below and confirmed that it resolved the issue with the two characters not approaching each other in our test. The revised wording makes both characters' walking movements and the decreasing distance between them explicit.
Please also remove this sentence from the original prompt:
"Preserve their faces, bodies, identities, clothing, screen positions across every hard cut."
The instruction to preserve screen positions may conflict with the requested movement. We have not isolated it as the sole cause, but removing it avoids that potential conflict.
Please try this action prompt:
Full body shot of two people on the Tatooine planet matching Image 5. Image 1 is on the left side with his orange lightsaber. Image 2 is on the right side with her green lightsaber.
Both characters actively walk toward each other at the same time. The man walks from the left toward the woman, while the woman walks from the right toward the man. They each take several clear, visible steps forward, steadily closing the distance between them until they stop face-to-face within lightsaber striking distance. Keep both characters' full bodies and walking movements visible throughout their approach.
Then the woman (Image 2) laughs and says, "Okay, let's start." The man (Image 1) grins and says, "Alright, if that's what you want." They both raise their lightsabers and clash the blades in the air, and the sound of clashing lightsabers fills the area.
Regarding the other issues you reported:
1.Character appearance: You mentioned that the scene matches Image 5, but the characters do not match Images 1 and 2. We did not observe significant character consistency issues in our tests. Please try again with the revised prompt.
2.Dialogue language: You reported speech in an unexpected language. Please paste the prompt into the appropriate prompt field again and retry.
3.Lightsabers: Please share the generated video so we can check whether the issue involves blade shape, color, handling, or the clash itself. In our tests, we noticed some occasional visual artifacts, but nothing unusual beyond those.
4.Model loading: You mentioned that loading stalled at “Requested to load LTXAV” despite having 128GB of unified memory, so you changed the diffusion model weight setting from default to fp8_e5m2. We cannot yet determine whether this change affected the generated results or what caused the loading stall.
If any of these issues remain after trying the revised prompt, please send your exported workflow JSON, the generated video, your hardware and operating system details, and the loading log. These will help us compare your setup with ours and investigate the remaining issues separately.
Thanks again for the detailed feedback.