LuxTTS / README.md
freeman412's picture
Mirror of YatharthS/LuxTTS (Apache-2.0) — pins a first-run dependency for gloam.fm
bad8917 verified
|
Raw
History Blame Contribute Delete
1.71 kB
metadata
license: apache-2.0
language:
  - en
pipeline_tag: text-to-speech

LuxTTS (mirror)

This is an unmodified mirror of YatharthS/LuxTTS. All credit for the model belongs to its author, YatharthS — original code and usage docs: https://github.com/ysharma3501/LuxTTS.

We mirror it only to pin a dependency: gloam.fm downloads this model during first-run setup, so an upstream rename or deletion would break setup for every new user. The weights here are byte-identical to upstream. If you are looking for LuxTTS itself, use the original repo — it is the canonical source and the place to star, file issues, and follow updates.

Released under Apache-2.0, same as upstream.


The upstream model card follows, unmodified.

LuxTTS

This is the model for LuxTTS, a lightweight zipvoice based text-to-speech model designed for high quality voice cloning and realistic generation at speeds exceeding 150x realtime.

Main features

  • Voice cloning: SOTA voice cloning on par with models 10x larger.
  • Clarity: Clear 48khz speech generation unlike most TTS models which are limited to 24khz.
  • Speed: Reaches speeds of 150x realtime on a single GPU and faster then realtime on CPU's as well.
  • Efficiency: Fits within 1gb vram meaning it can fit in any local gpu.

Details

  • Based on ZipVoice, distilled to 4steps.
  • Uses 48khz vocoder instead of 24khz vocoder.
  • Implemented higher quality sampling technique then standard euler.

Usage

Please check out the repo for usage: https://github.com/ysharma3501/LuxTTS.git

License

Model and code is released under Apache-2.0 license.