KoshiMazaki's picture
Update scripts/dataset/README.md
707dac7 verified
|
Raw
History Blame Contribute Delete
2.61 kB
# Dataset preparation
LTX's trainer takes a dataset; it does not help you *build* one. For an audio LoRA the dataset is the challange, the model learns a transformation, so
every item must be the same performance twice, once dry and once through a real
acoustic treatment, aligned and gain-matched. I also used two iterations of each dry sound source so model can learn on variety.
These are the scripts that built the AKUSPACE set (266 pairs, 14 sources Γ— 19 cells).
Set `AKUSPACE_DATASET` to your dataset root, or run from the repo and the
relative default applies.
| Script | Does |
|---|---|
| `build_v5_rooms.py` | dry/wet room pairs at three levels per room |
| `build_v5_outdoor.py` | outdoor beds; projects the dry out before rebuilding |
| `build_v5_sfx.py` | Eurorack modular-delay pairs |
| `build_eurorack_dry.py` | assembles the dry side fed through the hardware |
| `make_manifest_v5.py` | writes the training CSV with captions and splits |
| `verify_dataset.py` | checks every pair: length, rate, ceiling, alignment |
| `check_limiting.py` | catches limiting/clipping introduced by a render chain |
| `check_clips.py` | duration and sample-rate conformance |
| `wetdry.py` | measures the wet/dry relationship of a rendered pair |
| `slice_ableton_stems.py` | cuts long recordings into per-source clips |
## The rules these encode
Each one cost real time to learn.
**Never normalise dry and wet independently.** Apply the same gain to both, or
the pair no longer describes a transformation and the model learns loudness
instead of acoustics.
**Normalise by active loudness, not peak.** Sparse material β€” handclaps
especially β€” reads ~13 dB quiet on peak-based normalisation and ends up
mismatched against sustained sources.
**An ambience bed is not a reverb tail.** A tail is part of the signal and
scales cleanly; a bed is a separate recording, so boosting it amplifies its own
noise floor. Scale beds down only, and only a little. This is why the outdoor
cells ship two levels while rooms and SFX ship three.
**Check the dry fit before trusting any ambience measurement.** `R = wet βˆ’ dry`
assumes the render carries the dry at exactly 1.00Γ—. When it does not, a
negative ghost of the dry hides inside R and dominates every derived number β€”
which produced months of misleading ambience statistics here before
`verify_dataset.py` caught it. Confirm fit β‰ˆ 1.000 first.
**The trigger word belongs in preprocessing, not the manifest.** It is prepended
to training captions by the preprocessor and must appear in inference prompts;
storing it in the CSV double-applies it.