Instructions to use ModelsLab/midashenglm-gen-wer-lora with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ModelsLab/midashenglm-gen-wer-lora with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
File size: 7,804 Bytes
3d3c7b8 a98dc75 3d3c7b8 e9e75f4 3d3c7b8 e9e75f4 3d3c7b8 e9e75f4 3d3c7b8 e9e75f4 3d3c7b8 e9e75f4 3d3c7b8 e9e75f4 3d3c7b8 e9e75f4 3d3c7b8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 | ---
license: apache-2.0
base_model: mispeech/midashenglm-gen
library_name: peft
tags:
- audio-generation
- text-to-audio
- lora
- flow-matching
datasets:
- mythicinfinity/libritts_r
- mispeech/MECAT-Caption
---
# ModelsLab/midashenglm-gen-wer-lora
A LoRA adapter for [`mispeech/midashenglm-gen`](https://huggingface.co/mispeech/midashenglm-gen),
trained across all five of the model's capabilities β speech, sound effects,
music, ambience and mixed scenes β rather than trading one for another.
## Hear it
Same prompt, same seed in both arms β `generate()` seeds the global RNG, so the
solver draws identical noise and every audible difference is the adapter.
**Speech, 4-word line β "The river remembers everything." The bucket that went 13.9% β 0.0% WER; the base arm garbles it.**
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/speech-4w__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/speech-4w__V8.wav"></audio> |
**Speech, 16-word line β both arms intelligible; the adapter is cleaner and tighter.**
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/speech-16w__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/speech-16w__V8.wav"></audio> |
**SFX β footsteps on gravel. Texture and caption match improve; the adapter clip runs shorter, the one real remaining cost.**
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/sfx-gravel__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/sfx-gravel__V8.wav"></audio> |
**Music β slow solo piano. The adapter plays longer and closer to the caption at this seed.**
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/music-piano__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/music-piano__V8.wav"></audio> |
**Ambience β steady rain on a metal roof. Near-identical by design: this arm held.**
| base | scene-v8 |
|---|---|
| <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/ambience-rain__BASE.wav"></audio> | <audio controls src="https://huggingface.co/ModelsLab/midashenglm-gen-wer-lora/resolve/main/samples/ambience-rain__V8.wav"></audio> |
## Why
Two weak spots in the base model, measured rather than assumed.
**Short lines.** A four-word line scores around 42% WER where a sixteen-word one
scores near zero. The failures are onsets: the model needs a moment to settle
and a short line does not give it one. Prompt-side remedies make it worse, so it
is a property of the weights.
**Sound effects.** The model's worst benchmark β AudioCaps FAD 5.01 against
TangoFlux's 2.26 β which the paper attributes to training on mixed scenes rather
than dedicated sound-effect corpora.
## Results
Same prompt, same seed, base weights against these adapters. `generate()` seeds
the global RNG, so the flow-matching solver draws identical noise in both arms
and every difference is the adapters.
| capability | metric | base | tuned | change |
|---|---|---:|---:|---:|
| speech | WER | 5.6% | 0.2% | -5.4% |
| sfx | CLAP text | 0.4064 | 0.3837 | -0.0226 |
| sfx | CLAP real | 0.4152 | 0.4201 | +0.0049 |
| music | CLAP text | 0.4736 | 0.4673 | -0.0063 |
| music | CLAP real | 0.5928 | 0.6012 | +0.0084 |
| ambience | CLAP text | 0.2442 | 0.2070 | -0.0371 |
| ambience | CLAP real | 0.3667 | 0.3732 | +0.0065 |
| mixed | WER | 4.8% | 4.7% | -0.1% |
| mixed | CLAP text | 0.1708 | 0.1745 | +0.0037 |
| mixed | CLAP real | 0.7014 | 0.6943 | -0.0072 |
| speech | WER, 4-word lines | 13.9% | 0.0% | -13.9% |
| speech | WER, 9-word lines | 3.2% | 0.0% | -3.2% |
| speech | WER, 16-word lines | 1.9% | 0.6% | -1.2% |
| speech | WER, 27-word lines | 3.4% | 0.3% | -3.1% |
| speech | mean length | 4.73s | 4.47s | -0.26s |
| sfx | mean length | 8.08s | 6.93s | -1.15s |
| music | mean length | 9.42s | 9.19s | -0.23s |
| ambience | mean length | 10.43s | 10.29s | -0.14s |
| mixed | mean length | 4.97s | 4.47s | -0.50s |
CLAP text is "does it match the caption". CLAP real is cosine to the centroid of
genuine MECAT audio of that category β "does it sound like the real thing",
which text similarity alone cannot see. WER cannot see either: a clip with no
words in it scores 0% however badly the room tone came out.
**WER improved by 5.4 points, and that part is solid. The scene arms are not.**
Caption adherence or realism fell on: sfx clap_text -0.023, ambience clap_text -0.037.
**Use this for speech.** For beds, effects and mixed scenes, A/B it against the base model on your own prompts first β the held-out flow loss improved on every capability, and on the scene arms that did not translate.
## Training
| | |
|---|---|
| Method | LoRA r=32, alpha=64, on LLM attention/MLP and the flow-matching DiT |
| Trainable | 72876032 of 2.89 B |
| Precision | float32 with TF32 matmuls, matching how the base model is served |
| Optimiser | AdamW-8bit, lr 5e-05, cosine decay, grad clip 1.0 |
| Steps | 2500 at an effective batch of 16 |
| Capability mix | sfx 22%, ambience 20%, speech 20%, mixed 20%, music 18% |
| Speech data | LibriTTS-R + MECAT S00; clips under 3 s oversampled 2.5x |
| Scene data | MECAT-Caption, all eight categories |
| Anchor | vector field pulled towards the frozen base, weight 0.5 on 35% of batches |
Base-model hyperparameters follow arXiv:2608.11804 Β§3.3 where they apply
(lambda_stop = 0.01, grad clip 1.0).
### Avoiding catastrophic forgetting
Four things, because one is not enough:
1. **LoRA rather than a full fine-tune.** Rank 32 bounds how far the weights
can move at all.
2. **A capability-weighted mix.** The sampler draws by capability, not by pool.
At natural frequency the speech buckets are ~80% of the clips and the model
would simply get better at reading aloud.
3. **An anchor loss.** The tuned vector field is pulled towards the frozen base's
on identical noise and timesteps β a trust region around the base model,
stated directly rather than hoped for.
4. **A per-capability gate.** Held-out flow loss is tracked for each capability
separately, and this checkpoint was only saved because the mean improved *and*
no single capability regressed past 2%.
## Use
```python
from peft import PeftModel
from transformers import AutoModel
import soundfile as sf
model = AutoModel.from_pretrained("mispeech/midashenglm-gen", trust_remote_code=True)
model = PeftModel.from_pretrained(model, "ModelsLab/midashenglm-gen-wer-lora").merge_and_unload().cuda().eval()
result = model.generate(
"<|caption|> A close, clean recording of a single narrator, with nothing behind it. "
"<|asr|> The river remembers everything. "
"<|speech|> A woman in her forties, warm and even, speaking clearly. "
"<|sfx|> <|unknown|> <|music|> <|unknown|> <|env|> <|unknown|>",
seed=11,
)
sf.write("out.wav", result["audio"], result["sample_rate"])
```
Tag order matters more than anything else in the prompt: `<|asr|>` comes
**before** `<|speech|>`. Reversed, the model produces fluent unrelated speech β
14.2% mean WER against 373%.
## Licences
Apache 2.0, following the base model. Training data: LibriTTS-R (CC-BY-4.0) and
MECAT-Caption (CC-BY-3.0), both attribution-only.
Training code: [`audio-scenegen/training`](https://github.com/adhikjoshi/audio-scenegen).
|