File size: 5,665 Bytes
2771d54
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
---
library_name: diffusers
tags:
- modular-diffusers
- diffusers
- echo-wm
- text-to-image
---
This is a modular diffusion pipeline built with 🧨 Diffusers' modular pipeline framework.

**Pipeline Type**: EchoWMBlocks

**Description**: 

This pipeline uses a 5-block architecture that can be customized and extended.

## Example Usage

[TODO]

## Pipeline Architecture

This modular pipeline is composed of the following blocks:

1. **text** (`LTX2TextConditioningStep`)
   - Text-conditioning stage for LTX-2.X: encodes the prompt(s), then runs the text connectors to produce the video/audio-branch connector embeddings the denoiser consumes. Outputs stay at one row per prompt -- the denoise stage expands them by `num_videos_per_prompt` -- so they can be reused across denoise runs.
2. **camera** (`EchoWMCameraConditionStep`)
3. **image_encoder** (`EchoWMVaeEncoderStep`)
   - VAE encoder step that encodes the input `image` into normalized latents for image-to-video generation.
4. **denoise** (`EchoWMImage2VideoCoreDenoiseStep`)
   - Denoise block (image-to-video) that expands the text conditioning by `num_videos_per_prompt`, adds image conditioning and runs the joint denoising loop.
5. **decode** (`EchoWMDecoderStep`) 

## Model Components

1. text_encoder (`PreTrainedModel`)
2. tokenizer (`PreTrainedTokenizerBase`)
3. connectors (`LTX2TextConnectors`)
4. transformer (`EchoWMTransformer3DModel`)
5. vae (`AutoencoderKLLTX2Video`)
6. video_processor (`VideoProcessor`)
7. scheduler (`FlowMatchEulerDiscreteScheduler`)
8. audio_vae (`AutoencoderKLLTX2Audio`)
9. guider (`LTX2Guidance`)
10. audio_guider (`LTX2Guidance`)
11. vocoder (`LTX2Vocoder`) 

## Input/Output Specification

**Inputs:**

- `prompt` (`str`): The prompt or prompts to guide image generation.
- `negative_prompt` (`str`, *optional*): The prompt or prompts not to guide the image generation.
- `max_sequence_length` (`int`, *optional*, defaults to `1024`): Maximum sequence length for prompt encoding.
- `action` (`str`): WASD/IJKL action program.
- `height` (`int`, *optional*, defaults to `704`): The height in pixels of the generated image.
- `width` (`int`, *optional*, defaults to `1280`): The width in pixels of the generated image.
- `num_frames` (`int`, *optional*, defaults to `241`): Number of output video frames.
- `frame_rate` (`float`, *optional*, defaults to `24.0`): Output video frame rate.
- `translation_speed` (`float`, *optional*, defaults to `0.05`): Per-frame camera translation speed for W/A/S/D actions.
- `rotation_speed_deg` (`float`, *optional*, defaults to `0.5`): Per-frame camera yaw speed in degrees for J/L actions.
- `pitch_speed_deg` (`float`, *optional*, defaults to `0.2`): Per-frame camera pitch speed in degrees for I/K actions.
- `pitch_limit_deg` (`float`, *optional*, defaults to `60.0`): Maximum absolute camera pitch in degrees.
- `fov_deg` (`float`, *optional*, defaults to `70.0`): Horizontal camera field of view in degrees.
- `num_videos_per_prompt` (`int`, *optional*, defaults to `1`): The number of images to generate per prompt.
- `image` (`Image | list`): Reference image(s) for denoising. Can be a single image or list of images.
- `image_crf` (`int`, *optional*): H.264 CRF used to re-compress the conditioning `image` before VAE encode, matching the compression the model was trained against. `None` (default) resolves from the text-encoder generation (33 through LTX-2.3, 18 for LTX-2.5). Pass `0` to skip re-compression. Requires a `PIL.Image.Image` when re-compression runs.
- `generator` (`Generator`, *optional*): Torch generator for deterministic generation.
- `num_inference_steps` (`int`, *optional*, defaults to `30`): The number of denoising steps.
- `timesteps` (`Tensor`, *optional*): Timesteps for the denoising process.
- `sigmas` (`list`, *optional*): Custom sigmas for the denoising process.
- `latents` (`Tensor`, *optional*): Pre-generated noisy latents for image generation.
- `noise_scale` (`float`, *optional*): Interpolation factor between random noise and any provided latents. `None` (default) resolves to 0.0, which keeps the provided latents.
- `audio_latents` (`Tensor`, *optional*): Optional pre-encoded audio latents; random noise is used when not provided.
- `**denoiser_input_fields` (`None`, *optional*): conditional model inputs for the denoiser: e.g. prompt_embeds, negative_prompt_embeds, etc.
- `use_cross_timestep` (`bool`, *optional*, defaults to `True`): Whether to condition the transformer on a separate per-token cross timestep (LTX-2.3+).
- `attention_kwargs` (`dict`, *optional*): Additional kwargs for attention processors.
- `output_type` (`str`, *optional*, defaults to `pil`): Output format: 'pil', 'np', 'pt'.
- `decode_timestep` (`None`, *optional*, defaults to `0.0`): The timestep at which the VAE decodes the final latents.
- `decode_noise_scale` (`None`, *optional*): Noise interpolation factor applied to the latents at the decode timestep.
- `vae_tiling` (`bool`, *optional*, defaults to `True`): Enable spatial and temporal VAE decoding tiles to reduce peak memory usage.
- `vae_tile_size` (`int`, *optional*, defaults to `512`): Spatial tile long-side size in pixels; the short side follows the video aspect ratio.
- `vae_tile_overlap` (`int`, *optional*, defaults to `64`): Spatial tile overlap in pixels.
- `vae_temporal_tile_size` (`int`, *optional*, defaults to `64`): Temporal tile size in sample frames, excluding the causal boundary frame.
- `vae_temporal_tile_overlap` (`int`, *optional*, defaults to `24`): Temporal tile overlap in sample frames.

**Outputs:**

- `videos` (`list`): The generated videos.
- `audio` (`Tensor`): The generated audio waveform.