VoiceEraser ๐Ÿ”‡ โ€” audio demo

Speaker unlearning for zero-shot TTS. One host-independent edit (cond-DIM) makes a released model stop cloning a chosen voice, while everyone else is untouched. Listen below.

What you are hearing, for each forget speaker (the voice being opted out):

1 ยท Erasure across three backbones

Same forget speakers, each cloned then unlearned on three architecturally different autoregressive backbones. Left = baseline clone, right = after unlearning.

Forget speakerXTTS-v2Tortoise-TTSIndexTTS-1.5

2 ยท Relearn stress-test (XTTS-v2)

An adversary who downloads the edited open weights and re-finetunes them to recover the voice. Recovery is only partial โ€” the whole point of the paper's honest robustness result.

Forget speakerBaseline cloneAfter unlearningAfter relearn attack