File size: 1,714 Bytes
bad8917
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
---
license: apache-2.0
language:
- en
pipeline_tag: text-to-speech
---

# LuxTTS (mirror)

**This is an unmodified mirror of [`YatharthS/LuxTTS`](https://huggingface.co/YatharthS/LuxTTS).**
All credit for the model belongs to its author, [YatharthS](https://huggingface.co/YatharthS)
— original code and usage docs: <https://github.com/ysharma3501/LuxTTS>.

We mirror it only to pin a dependency: [gloam.fm](https://gloam.fm) downloads this model
during first-run setup, so an upstream rename or deletion would break setup for every new
user. The weights here are byte-identical to upstream. If you are looking for LuxTTS itself,
**use the original repo** — it is the canonical source and the place to star, file issues,
and follow updates.

Released under Apache-2.0, same as upstream.

---

_The upstream model card follows, unmodified._

## LuxTTS

This is the model for LuxTTS, a lightweight zipvoice based text-to-speech model designed for
high quality voice cloning and realistic generation at speeds exceeding 150x realtime.

### Main features
- Voice cloning: SOTA voice cloning on par with models 10x larger.
- Clarity: Clear 48khz speech generation unlike most TTS models which are limited to 24khz.
- Speed: Reaches speeds of 150x realtime on a single GPU and faster then realtime on CPU's as well.
- Efficiency: Fits within 1gb vram meaning it can fit in any local gpu.

### Details
- Based on ZipVoice, distilled to 4steps.
- Uses 48khz vocoder instead of 24khz vocoder.
- Implemented higher quality sampling technique then standard euler.

### Usage
Please check out the repo for usage: https://github.com/ysharma3501/LuxTTS.git

### License
Model and code is released under Apache-2.0 license.