akuspace-ltx25 / scripts /README.md
KoshiMazaki's picture
scripts: add README.md
971526d verified
|
Raw
History Blame Contribute Delete
4.02 kB
# Scripts
The tooling that produced this LoRA and its examples, so the results can be
reproduced and checked rather than taken on trust.
## Inference
**`ltx_a2a_generate_multi.py`** β€” audio-to-audio generation from a TSV manifest.
The LTX trainer ships no standalone a2a inference CLI; `validation_runner.py` is
that code path, so this drives `scripts/train.py` with `optimization.steps: 1`
and `checkpoints.no_resume: true` and collects what validation writes.
Its reason for existing: the stock helper pins one `--reference` to every
prompt, so an N-source sweep costs N model loads per checkpoint. Validation
samples each carry their own `conditions[].audio`, so one load can cover every
(source, prompt) pair. A 3-source Γ— 19-cell sweep drops from 3 model loads to 1,
and loading dominates runtime.
```
ltx_a2a_generate_multi.py --checkpoint <lora.safetensors> \
--base-config <inference.yaml> --manifest sweep.tsv --output-dir out/
```
Manifest is `label<TAB>reference_wav<TAB>prompt`. The trigger word is prepended
by the tool, matching how training captions were built β€” never put it in the
manifest.
**`build_sweep_manifest.py`** β€” builds that manifest from the trained caption
vocabulary. Captions are assembled from the exact strings in the training
manifest with the source type substituted, so prompts stay in distribution.
Cell sets: `sweetspot`, `compare`, `all` (12 room cells), `placesfx` (outdoor +
SFX), `full` (all 19).
**`build_heldout_refs.py`** β€” rebuilds the 15.08 s held-out references from the
dataset's dry sources. Every source is exactly 6.000 s and the render length is
377 frames at 25 fps, so the reference is the dry clip repeated with a gap and
trimmed. No dependencies beyond numpy.
## Verification
**`measure_timing.py`** β€” envelope cross-correlation between a dry reference and
its render, in 2 ms RMS windows over Β±250 ms. This reproduces the timing table
in the model card, so those numbers can be checked rather than believed.
Read the correlation column carefully: it measures how much the envelope
*changed*, and reverb changes it by design β€” most on transient-dense material
where the tail fills the gaps between hits. Low r on percussion is the effect
working. It is a timing measure, not a quality measure.
## Training / preprocessing
**`box_b_trainer_setup.sh`** β€” dataset preprocess plus the **mandatory flatten
step**.
`process_dataset.py` mirrors the dataset-JSON relative paths, writing
`audio_latents/audio/targets/X.pt` and
`reference_audio_latents/audio/REFERENCES/X.pt`. But `datasets.py::_discover_samples`
takes the primary source's `rel_path` and requires that identical path in every
other source, so targets and references never match. The run then reports
**"No valid samples found"** with per-source counts of zero while all three
directories are visibly full. Flattening all three to a single level fixes it.
This is silent and will hit anyone training a multi-source a2a IC-LoRA.
Also note `keep_last_n` is a rolling window β€” set it to `-1` to keep every
checkpoint, or early ones are deleted without warning.
Inference needs preprocessed latents too, because `train.py` builds a
dataloader before validation fires. It does not need *all* of them: a handful of
pairs satisfies it in under a minute.
## Video
**`build_flf25.py`** β€” first-and-last-frame LTX-2.5 graphs in ComfyUI API
format, both frames via chained `LTXVAddGuide` at `frame_idx` 0 and βˆ’1. Two
presets, `distilled` and `dev`; see `WORKING_CONFIG.md`. Every node class and
input name is validated against a live `object_info` dump before the graph is
written.
**`WORKING_CONFIG.md`** β€” the verified-good settings, and the trap that the two
transformers need opposite sampling regimes.
Note the a2a LoRA does **not** belong inside a video graph: it is
reference-conditioned, so with no reference it is off-distribution and yields
near-silence. Render video, take its audio, run the a2a pass, remux. That also
gives the base-model baseline for free.