kgrosero14 commited on
Commit
76e59ef
·
verified ·
1 Parent(s): 40b5ef7

Update model card: current presets, venue, and usage snippet

Browse files

Replaces the five-preset table with the two presets that ship (default, torgo) and fixes the usage snippet, which called load_preset("star_v5") for a preset that no longer exists. Adds the GPU memory the training recipe needs and the SLT 2026 venue; drops internal run filenames.

Files changed (1) hide show
  1. README.md +53 -49
README.md CHANGED
@@ -14,14 +14,14 @@ tags:
14
 
15
  # ChiReSSD
16
 
17
- Speaker-preserving reconstruction of pediatric disordered speech.
18
 
19
- Given a target transcript and a short reference recording from a child with a speech sound
20
- disorder (SSD), ChiReSSD synthesizes the target utterance with canonical pronunciation while
21
- keeping the child's voice and prosody. Pronunciation enters through the text pathway; identity
22
- and prosody come from the style pathway. That separation is the point: ordinary style-preserving
23
- TTS treats disordered articulation as part of the speaker's style and so reproduces the
24
- mispronunciation it was meant to correct.
25
 
26
  - Code: <https://github.com/Lab-MSP/ChiReSSD>
27
  - Paper: *Generative Reconstruction of Pediatric Disordered Speech for Automated Clinical
@@ -29,27 +29,28 @@ mispronunciation it was meant to correct.
29
 
30
  ## Model description
31
 
32
- StyleTTS2 fine-tuned on child disordered speech. Two 128-dimensional style vectors are extracted
33
- from the reference recording — acoustic (timbre) and prosodic — and interpolated with a style
34
- sampled from the adapted diffusion prior. `alpha` weights the acoustic side and `beta` the
35
- prosodic; 1.0 is fully diffusion-sampled, 0.0 fully reference-driven.
36
 
37
- It works on unseen speakers from a reference as short as a few seconds. No per-child model is
38
- trained, and no paired typical/atypical recordings are required.
39
 
40
  ## Base model and license chain
41
 
42
  Fine-tuned from the StyleTTS2 LibriTTS second-stage checkpoint by
43
- [yl4579](https://github.com/yl4579/StyleTTS2) (MIT). **That base checkpoint is not redistributed
44
- here** — obtain it from the upstream release. The frozen helper models (ASR text aligner, JDCNet
45
- pitch extractor, PL-BERT) likewise ship inside the upstream repository and are not redistributed.
 
46
 
47
  ChiReSSD's own modification to StyleTTS2 is two lines, published as patch files in the code
48
  repository rather than as a fork.
49
 
50
  ## Intended use
51
 
52
- Research on speech reconstruction and on automated clinical evaluation of pediatric speech sound
53
  disorders.
54
 
55
  ## Out of scope
@@ -64,9 +65,9 @@ disorders.
64
 
65
  Child speech from UltraSuite/CLP-derived recordings, obtained under their own data use
66
  agreements. **The corpus is not released here and is not redistributable**: it is identifiable
67
- child clinical speech. See `DATA.md` in the code repository for how to obtain comparable data;
68
- <https://huggingface.co/datasets/changelinglab/ultrasuite-benchmark> is a suitable starting point
69
- for fine-tuning your own model.
70
 
71
  ## Training configuration
72
 
@@ -84,41 +85,41 @@ for fine-tuning your own model.
84
  | Sample rate | 24 kHz |
85
  | LR schedule | OneCycleLR, `pct_start=0.1` (upstream: 0) |
86
 
87
- The heavy pitch weighting is deliberate: the pitch extractor was pretrained on adult voices, and
88
- children's F0 is both higher and more variable, so it needs the strongest adaptation of any
89
- component. Conversely, only four epochs — longer schedules start fitting the disordered
90
  articulation itself.
91
 
 
 
 
92
  ## Inference
93
 
 
 
94
  | Preset | alpha | beta | steps | Purpose |
95
  |---|---|---|---|---|
96
- | `star_v5` | 0.8 | 0.5 | 10 | Released operating point |
97
- | `star_v5_alt` | 1.0 | 0.4 | 10 | Variant |
98
- | `star_oneshot_base` | 1.0 | 0.5 | 5 | One-shot baseline (no fine-tuning) |
99
- | `star_adult_tts` | 0.3 | 0.7 | 5 | Standard adult-TTS baseline |
100
- | `torgo_v5` | 1.0 | 0.5 | 15 | Adult dysarthric speech |
101
 
102
  `alpha` is high on purpose. A low `alpha` leans on the reference acoustics, which is exactly
103
  where the disordered articulation lives.
104
 
105
- **Synthesis is stochastic.** The initial style latent is zeros rather than a Gaussian draw, but
106
- the ADPM2 sampler is ancestral and adds fresh noise at every step, so repeated calls differ in
107
- waveform and in duration. Pass a `seed` for reproducible output.
108
 
109
  ## Evaluation
110
 
111
- See the paper. **This model repository does not reproduce the paper's reported numbers**, and the
112
- code repository does not either: the evaluation corpora are not redistributable, so the numbers
113
- cannot be independently recomputed from what is published here.
114
 
115
  ## Ethical considerations
116
 
117
  This model is trained on identifiable clinical recordings of children and is designed to
118
- reproduce a child's vocal identity accurately. High speaker similarity is the method's goal and
119
- therefore also its misuse vector: the same property that makes feedback usable in a child's own
120
- voice makes the model a voice-cloning tool for a minor. Use it only with appropriate consent and
121
- ethical oversight.
122
 
123
  ## Usage
124
 
@@ -127,31 +128,34 @@ from chiressd.model import ChiReSSD
127
  from chiressd.config import load_preset
128
 
129
  model = ChiReSSD.from_pretrained() # downloads this checkpoint
130
- style = model.compute_style("child_reference.wav")
131
  wav = model.synthesize(
132
- "butterfly butterfly butterfly", style, seed=1234, **load_preset("star_v5").as_kwargs()
 
 
 
133
  )
134
  ```
135
 
136
- Run `chiressd-setup` first: it clones and patches the upstream StyleTTS2 checkout that supplies
137
- the architecture and frozen helper models.
138
 
139
  ## Checkpoint provenance
140
 
141
- Derived from training checkpoint `epoch_2nd_00003.pth` of the v5 run by keeping `state['net']`
142
- only, removing the `module.` prefix that DataParallel added to ten of the thirteen submodules,
143
- and making every tensor detached, CPU-resident and contiguous. Precision is unchanged (float32;
144
- no fp16 cast, which would alter outputs).
145
 
146
  - 2258 MB → 767 MB; the 1491 MB removed is optimizer state, which inference never reads
147
  - `sha256` of `model.pth`: `6c5faf5b4967b4d26cb4c58d4517dcc477729a579039316741bd88b9bd0f16e0`
148
  - Verified bit-exact against the training checkpoint it was stripped from: under a fixed seed,
149
- the two produce an identical 256-d style vector and identical waveform samples, so the strip
150
- provably changes no output
151
 
152
  Full details in `manifest.json`.
153
 
154
  ## Citation
155
 
156
  Rosero, Yeo, Mortensen, Van't Slot, Hallac, and Busso. *Generative Reconstruction of Pediatric
157
- Disordered Speech for Automated Clinical Evaluation.*
 
 
14
 
15
  # ChiReSSD
16
 
17
+ Speaker-preserving reconstruction of disordered speech.
18
 
19
+ Given a target transcript and a short reference recording of a speaker, ChiReSSD synthesizes
20
+ that transcript with canonical pronunciation while keeping the speaker's voice and prosody.
21
+ Pronunciation enters through the text pathway; identity and prosody come from the style
22
+ pathway. That separation is the point: ordinary style-preserving TTS treats disordered
23
+ articulation as part of the speaker's style and so reproduces the mispronunciation it was
24
+ meant to correct.
25
 
26
  - Code: <https://github.com/Lab-MSP/ChiReSSD>
27
  - Paper: *Generative Reconstruction of Pediatric Disordered Speech for Automated Clinical
 
29
 
30
  ## Model description
31
 
32
+ StyleTTS2 fine-tuned on child disordered speech. Two 128-dimensional style vectors are
33
+ extracted from the reference recording — acoustic (timbre) and prosodic — and interpolated
34
+ with a style sampled from the adapted diffusion prior. `alpha` weights the acoustic side and
35
+ `beta` the prosodic; 1.0 is fully diffusion-sampled, 0.0 fully reference-driven.
36
 
37
+ It works on unseen speakers from a reference as short as a few seconds. No per-speaker model
38
+ is trained, and no paired typical/atypical recordings are required.
39
 
40
  ## Base model and license chain
41
 
42
  Fine-tuned from the StyleTTS2 LibriTTS second-stage checkpoint by
43
+ [yl4579](https://github.com/yl4579/StyleTTS2) (MIT). **That base checkpoint is not
44
+ redistributed here** — obtain it from the upstream release. The frozen helper models (ASR text
45
+ aligner, JDCNet pitch extractor, PL-BERT) likewise ship inside the upstream repository and are
46
+ not redistributed.
47
 
48
  ChiReSSD's own modification to StyleTTS2 is two lines, published as patch files in the code
49
  repository rather than as a fork.
50
 
51
  ## Intended use
52
 
53
+ Research on speech reconstruction and on automated clinical evaluation of speech sound
54
  disorders.
55
 
56
  ## Out of scope
 
65
 
66
  Child speech from UltraSuite/CLP-derived recordings, obtained under their own data use
67
  agreements. **The corpus is not released here and is not redistributable**: it is identifiable
68
+ child clinical speech. To fine-tune your own model,
69
+ <https://huggingface.co/datasets/changelinglab/ultrasuite-benchmark> is a suitable starting
70
+ point; see `DATA.md` in the code repository.
71
 
72
  ## Training configuration
73
 
 
85
  | Sample rate | 24 kHz |
86
  | LR schedule | OneCycleLR, `pct_start=0.1` (upstream: 0) |
87
 
88
+ The heavy pitch weighting is deliberate: the pitch extractor was pretrained on adult voices,
89
+ and children's F0 is both higher and more variable, so it needs the strongest adaptation of
90
+ any component. Conversely, only four epochs — longer schedules start fitting the disordered
91
  articulation itself.
92
 
93
+ Trained on 2× 48 GB GPUs. At batch 4 and `max_len` 600 the recipe needs more than 48 GB, so a
94
+ single smaller card requires lowering both.
95
+
96
  ## Inference
97
 
98
+ Two presets ship with the code:
99
+
100
  | Preset | alpha | beta | steps | Purpose |
101
  |---|---|---|---|---|
102
+ | `default` | 0.8 | 0.5 | 10 | The released operating point |
103
+ | `torgo` | 1.0 | 0.5 | 15 | Adult dysarthric speech |
 
 
 
104
 
105
  `alpha` is high on purpose. A low `alpha` leans on the reference acoustics, which is exactly
106
  where the disordered articulation lives.
107
 
108
+ **Synthesis is stochastic.** The initial style latent is zeros rather than a Gaussian draw,
109
+ but the ADPM2 sampler is ancestral and adds fresh noise at every step, so repeated calls
110
+ differ in waveform and in duration. Pass a `seed` for reproducible output.
111
 
112
  ## Evaluation
113
 
114
+ See the paper.
 
 
115
 
116
  ## Ethical considerations
117
 
118
  This model is trained on identifiable clinical recordings of children and is designed to
119
+ reproduce a speaker's vocal identity accurately. High speaker similarity is the method's goal
120
+ and therefore also its misuse vector: the same property that makes feedback usable in a
121
+ child's own voice makes the model a voice-cloning tool for a minor. Use it only with
122
+ appropriate consent and ethical oversight.
123
 
124
  ## Usage
125
 
 
128
  from chiressd.config import load_preset
129
 
130
  model = ChiReSSD.from_pretrained() # downloads this checkpoint
131
+ style = model.compute_style("speaker_reference.wav")
132
  wav = model.synthesize(
133
+ "butterfly butterfly butterfly",
134
+ style,
135
+ seed=1234,
136
+ **load_preset("default").as_kwargs(),
137
  )
138
  ```
139
 
140
+ Run `chiressd-setup` first: it clones and patches the upstream StyleTTS2 checkout that
141
+ supplies the architecture and frozen helper models.
142
 
143
  ## Checkpoint provenance
144
 
145
+ Derived from the fine-tuning run's final checkpoint by keeping `state['net']` only, removing
146
+ the `module.` prefix that DataParallel added to ten of the thirteen submodules, and making
147
+ every tensor detached, CPU-resident and contiguous. Precision is unchanged (float32; no fp16
148
+ cast, which would alter outputs).
149
 
150
  - 2258 MB → 767 MB; the 1491 MB removed is optimizer state, which inference never reads
151
  - `sha256` of `model.pth`: `6c5faf5b4967b4d26cb4c58d4517dcc477729a579039316741bd88b9bd0f16e0`
152
  - Verified bit-exact against the training checkpoint it was stripped from: under a fixed seed,
153
+ the two produce an identical 256-d style vector and identical waveform samples
 
154
 
155
  Full details in `manifest.json`.
156
 
157
  ## Citation
158
 
159
  Rosero, Yeo, Mortensen, Van't Slot, Hallac, and Busso. *Generative Reconstruction of Pediatric
160
+ Disordered Speech for Automated Clinical Evaluation.* In Proceedings of the IEEE Spoken
161
+ Language Technology Workshop 2026 (SLT '26), Palermo, Italy, 2026.