molm-audio / README.md
DaniilOr's picture
Update README; add legacy diffwave_v5_2step + vibevoice_14bit_v2 eval JSONs
d989fa7 verified
|
Raw
History Blame
6.09 kB
# MOLM-Audio: SPDMark-Style Segment-Wise Audio Watermarking
LoRA-routing audio watermarks for three generators (HiFi-GAN, VibeVoice acoustic
decoder, DiffWave). Per SPDMark (https://arxiv.org/abs/2512.12090), each audio is
split into S=8 segments; each segment carries an HMAC-derived M-bit message
embedded via a parallel LoRA "basis dictionary." The verifier matches recovered
per-segment bits against the expected HMAC sequence using Hungarian assignment +
Binomial hypothesis test.
## Repo layout
```
checkpoints/
diffwave_spdmark_spec/final/{lora_weights.pt, extractor.pt, diffwave_full.pt}
diffwave_v5_2step/final/{lora_weights.pt, extractor.pt, diffwave_full.pt}
hifigan_spdmark_spec/final/{lora_weights.pt, extractor.pt, hifigan_full.pt}
vibevoice_spdmark_spec/final/{lora_weights.pt, extractor.pt, model_full.pt}
vibevoice_14bit_v2/final/{lora_weights.pt, extractor.pt, model_full.pt}
eval/<run>_<regime>/results.json
README.md
```
`lora_weights.pt` + the base model from torchaudio/pretrained is sufficient for
inference. `*_full.pt` is provided for one-step loading where the base model
isn't available locally.
Training/test wavs live in **MOLM-Audio/molm-audio-data** under `data/`.
## Checkpoints
### Current SPDMark runs (spec-trained, with Hungarian verifier)
| Run | Backbone | Routing | Paths | Bits/seg | Train attack |
|---|---|---|---|---|---|
| `hifigan_spdmark_spec` | HiFi-GAN V3 (torchaudio LJSpeech) | `0,1,2,3,4,5,6` | 4 | 14 | `nvlceqr` |
| `vibevoice_spdmark_spec` | VibeVoice acoustic decoder | `0,1,2,3,4,5,6` (×2 slots) | 2 | 14 | `nvlceqr` |
| `diffwave_spdmark_spec` | DiffWave (LJSpeech) | `0,4,8,12,16,20,24` | 4 | 14 | `nvlceqr` |
### Legacy SPDMark runs (older eval pipeline; included for reference)
| Run | Backbone | Routing | Paths | Bits/seg | Train attack | Notes |
|---|---|---|---|---|---|---|
| `vibevoice_14bit_v2` | VibeVoice acoustic decoder | `0,1,2,3,4,5,6` (×2 slots) | 2 | 14 | `nvlceq` | Pre-SPDMark-temporal-attacks training; eval JSON has no Hungarian verify block. |
| `diffwave_v5_2step` | DiffWave (LJSpeech) | `0,4,8,12,16,20,24` | 2 | 7 | `nvlceq` | 2-step diffusion (`--diffusion_steps 2`), `lambda_perc 0.1` → louder watermark, faster inference. |
Spectral attack codes: `n` noise, `v` gain, `l` lowpass, `c` crop, `e` erase, `q` quantize, `r` resample. SPDMark temporal codes: `d` segment-drop, `s` segment-swap, `i` segment-insert.
## Inference
All three generators share a common eval interface; `--verify` enables the
Hungarian + Binomial verifier (auto-sets `--message_scheme hmac`).
### HiFi-GAN
```bash
python generate_molm_hifigan.py \
--lora_weights checkpoints/hifigan_spdmark_spec/final/lora_weights.pt \
--extractor_weights checkpoints/hifigan_spdmark_spec/final/extractor.pt \
--routing_blocks 0,1,2,3,4,5,6 --num_paths 4 --lora_rank 64 \
--test_dir <path-to-wavs> --num_samples 20 \
--chunked_generation --num_chunks 8 \
--verify --gamma_f 0.01 --gamma_v 0.01 --attacks nvlceqr \
--output_dir eval_out/hifigan_spec_spectral --device cuda
```
### DiffWave (current `spdmark_spec`)
```bash
python generate_molm_audio.py \
--diffwave_checkpoint pretrained/diffwave-ljspeech.pt \
--lora_weights checkpoints/diffwave_spdmark_spec/final/lora_weights.pt \
--extractor_weights checkpoints/diffwave_spdmark_spec/final/extractor.pt \
--routing_layers 0,4,8,12,16,20,24 --num_paths 4 --lora_rank 64 \
--diffusion_steps 4 \
--test_dir <path-to-wavs> --num_samples 20 \
--chunked_generation --num_chunks 8 \
--verify --gamma_f 0.01 --gamma_v 0.01 --attacks nvlceqr \
--output_dir eval_out/diffwave_spec_spectral --device cuda
```
### DiffWave (legacy `v5_2step`)
```bash
python generate_molm_audio.py \
--diffwave_checkpoint pretrained/diffwave-ljspeech.pt \
--lora_weights checkpoints/diffwave_v5_2step/final/lora_weights.pt \
--extractor_weights checkpoints/diffwave_v5_2step/final/extractor.pt \
--routing_layers 0,4,8,12,16,20,24 --num_paths 2 --lora_rank 64 \
--diffusion_steps 2 \
--test_dir <path-to-wavs> --num_samples 20 \
--chunked_generation --num_chunks 8 \
--verify --gamma_f 0.01 --gamma_v 0.01 --attacks nvlceqr \
--output_dir eval_out/diffwave_v5_2step_spectral --device cuda
```
Note: `lambda_perc 0.1` was used in training, so the watermark is louder
(higher bit accuracy, lower SNR) than `_spec`.
### VibeVoice (current `spdmark_spec`)
```bash
python generate_molm_vibevoice.py \
--components_dir pretrained/vibevoice \
--lora_weights checkpoints/vibevoice_spdmark_spec/final/lora_weights.pt \
--extractor_weights checkpoints/vibevoice_spdmark_spec/final/extractor.pt \
--routing_blocks 0,1,2,3,4,5,6 --num_paths 2 --lora_rank 16 --lora_alpha 8 \
--num_frames 8 \
--data_dir <path-to-precomputed-latents> --num_samples 20 \
--chunked_generation --num_chunks 8 \
--verify --gamma_f 0.01 --gamma_v 0.01 --attacks nvlceqr \
--output_dir eval_out/vibevoice_spec_spectral --device cuda
```
### VibeVoice (legacy `14bit_v2`)
Same flags as `spdmark_spec` but point at `checkpoints/vibevoice_14bit_v2/final/`.
Note: this run was trained without SPDMark temporal attacks (`dsi`), so its
robustness profile under `--attacks d`/`s`/`i` reflects the verifier alone.
## Eval snapshots
Each `eval/<run>_<regime>/results.json` is the literal output of the
corresponding `generate_molm_*.py --output_dir` call. Schemas differ between
the "current" runs (which include a `spdmark_verify` block with `tau_f`,
`tau_v`, `detection_rate_*`, `avg_|Q|_*`, `avg_bit_accuracy_on_Q_*`) and the
"legacy" runs (older schema, just raw bit accuracy and quality metrics).
DiffWave numbers under `eval/diffwave_spdmark_spec_step7000_*` are from an
undertrained snapshot (step 7000 of a 40k-step run) — re-evaluate against
`checkpoints/diffwave_spdmark_spec/final/` for the converged numbers.
## License / citation
LJSpeech audio is CC0. HiFi-GAN/DiffWave/VibeVoice base models retain upstream
licenses. Watermarking technique: Fares et al., SPDMark, arXiv 2512.12090.