Speaker unlearning for zero-shot TTS. One host-independent edit (cond-DIM) makes a released model stop cloning a chosen voice, while everyone else is untouched. Listen below.
cB) โ the released model cloning the target voice. It sounds like them.cE) โ after the cond-DIM edit: the identity is redirected away, so a
verifier no longer accepts it, yet the speech stays intelligible.cRL) โ a white-box adversary re-finetunes the released weights to
bring the voice back (only partially successful; XTTS-v2).Same forget speakers, each cloned then unlearned on three architecturally different autoregressive backbones. Left = baseline clone, right = after unlearning.
| Forget speaker | XTTS-v2 | Tortoise-TTS | IndexTTS-1.5 |
|---|
An adversary who downloads the edited open weights and re-finetunes them to recover the voice. Recovery is only partial โ the whole point of the paper's honest robustness result.
| Forget speaker | Baseline clone | After unlearning | After relearn attack |
|---|