voice-isolation-live / docs /step-1-demucs-setup.md
TonyLikeDev's picture
initial commit
0f623ae
|
Raw
History Blame Contribute Delete
1.25 kB

Step 1 β€” Demucs Setup

Status: DONE

What we did

Installed Demucs and ran it on a local audio file to confirm the model works before building anything around it.


What was installed

# CPU-only PyTorch (saves ~2GB vs CUDA build β€” no GPU on this machine)
pip install torch torchaudio --index-url https://download.pytorch.org/whl/cpu

# Demucs model
pip install demucs

# Required for saving output audio files
pip install torchcodec

How to run it manually

source venv/bin/activate
demucs --two-stems=vocals -n mdx_extra "path/to/audio.mp3"
  • -n mdx_extra β€” lighter model, faster on CPU
  • --two-stems=vocals β€” splits into two files: voice and background

Output

Results land in:

separated/mdx_extra/<track-name>/
β”œβ”€β”€ vocals.wav       # isolated voice
└── no_vocals.wav    # background / everything else

Notes

  • Processing time: ~40 seconds for a 33-second clip on ThinkPad X270
  • Fan noise mixed into voice is not fully removed β€” Demucs is a music separator, not a speech denoiser
  • Future fix: swap Demucs for DeepFilterNet (better for speech + background noise)
  • Model weights cached at ~/.cache/torch/hub/checkpoints/ β€” no re-download on next run