multimodalart HF Staff commited on
Commit
b4c6b95
·
verified ·
1 Parent(s): a42dbb8

Real speech for the audio example: LibriSpeech dev-clean 1462-170145-0022

Browse files
Files changed (1) hide show
  1. README.md +8 -0
README.md CHANGED
@@ -53,6 +53,14 @@ Rules the model imposes, enforced here before anything is uploaded:
53
  which is why the duration slider disappears when a single reference can set it, and comes back when two can or
54
  when the one that could is out of range.
55
 
 
 
 
 
 
 
 
 
56
  ## How the split is expressed
57
 
58
  `MiniMaxH3Ref2VABlocks` is a `SequentialPipelineBlocks` of eight steps:
 
53
  which is why the duration slider disappears when a single reference can set it, and comes back when two can or
54
  when the one that could is out of range.
55
 
56
+ ## Example assets
57
+
58
+ `examples/subject.png` is generated (FLUX.1-schnell), `examples/motion.mp4` is a synthetic clip from the parity
59
+ fixtures, and `examples/voice.wav` is utterance `1462-170145-0022` of
60
+ [LibriSpeech](https://www.openslr.org/12) `dev-clean` — CC BY 4.0, read from a public-domain LibriVox recording. It
61
+ is 16 kHz mono on purpose: the audio VAE wants 32 kHz, so the example exercises the `torchaudio` resample the
62
+ `ref2va` path needs.
63
+
64
  ## How the split is expressed
65
 
66
  `MiniMaxH3Ref2VABlocks` is a `SequentialPipelineBlocks` of eight steps: