Mirror of YatharthS/LuxTTS (Apache-2.0) — pins a first-run dependency for gloam.fm
bad8917 verified | license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: text-to-speech | |
| # LuxTTS (mirror) | |
| **This is an unmodified mirror of [`YatharthS/LuxTTS`](https://huggingface.co/YatharthS/LuxTTS).** | |
| All credit for the model belongs to its author, [YatharthS](https://huggingface.co/YatharthS) | |
| — original code and usage docs: <https://github.com/ysharma3501/LuxTTS>. | |
| We mirror it only to pin a dependency: [gloam.fm](https://gloam.fm) downloads this model | |
| during first-run setup, so an upstream rename or deletion would break setup for every new | |
| user. The weights here are byte-identical to upstream. If you are looking for LuxTTS itself, | |
| **use the original repo** — it is the canonical source and the place to star, file issues, | |
| and follow updates. | |
| Released under Apache-2.0, same as upstream. | |
| --- | |
| _The upstream model card follows, unmodified._ | |
| ## LuxTTS | |
| This is the model for LuxTTS, a lightweight zipvoice based text-to-speech model designed for | |
| high quality voice cloning and realistic generation at speeds exceeding 150x realtime. | |
| ### Main features | |
| - Voice cloning: SOTA voice cloning on par with models 10x larger. | |
| - Clarity: Clear 48khz speech generation unlike most TTS models which are limited to 24khz. | |
| - Speed: Reaches speeds of 150x realtime on a single GPU and faster then realtime on CPU's as well. | |
| - Efficiency: Fits within 1gb vram meaning it can fit in any local gpu. | |
| ### Details | |
| - Based on ZipVoice, distilled to 4steps. | |
| - Uses 48khz vocoder instead of 24khz vocoder. | |
| - Implemented higher quality sampling technique then standard euler. | |
| ### Usage | |
| Please check out the repo for usage: https://github.com/ysharma3501/LuxTTS.git | |
| ### License | |
| Model and code is released under Apache-2.0 license. | |