File size: 3,082 Bytes
01e924e
5c96cc1
a5a266c
 
 
 
 
 
 
 
5c96cc1
01e924e
5c96cc1
 
 
 
 
 
 
 
 
a5a266c
5c96cc1
 
 
 
 
a5a266c
5c96cc1
a5a266c
5c96cc1
a5a266c
5c96cc1
a5a266c
5c96cc1
 
 
 
 
 
 
 
 
 
 
a5a266c
 
 
 
 
 
 
 
 
 
 
 
 
 
5c96cc1
 
 
a5a266c
 
 
 
 
 
5c96cc1
a5a266c
 
5c96cc1
a5a266c
 
 
 
 
 
5c96cc1
a5a266c
5c96cc1
 
 
 
 
 
a5a266c
 
 
 
 
 
 
 
 
5c96cc1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a5a266c
5c96cc1
a5a266c
5c96cc1
a5a266c
5c96cc1
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
---

title: FroxAI Flex-Audio
emoji: 🦊
colorFrom: purple
colorTo: pink
sdk: gradio
sdk_version: 5.9.1
app_file: app.py
pinned: false
pipeline_tag: text-to-speech
license: other
license_name: proprietary

🦊 FroxAI β€” Flex-Audio

Fine-tuned "XTTS-v2" (https://huggingface.co/coqui/XTTS-v2) text-to-speech model with real voice cloning, built by FroxAI.

πŸŽ™οΈ Demo

The repository includes a Gradio application in "app.py" with:

- Text input
- Language selection
- Voice selection
- Generated WAV output
- Clickable demo examples

Run the demo

python app.py

Β«Hugging Face: A model repository does not automatically execute "app.py" on its main page. The interactive model-page widget is available only when the model/task is supported by an Inference Provider. For this custom Gradio interface, deploy "app.py" as a Hugging Face Space to let visitors test it directly in a browser.Β»

What's included

- "model.pth" β€” fine-tuned model weights
- "config.json" β€” base XTTS-v2 configuration required to reload the model
- "vocab.json" β€” base model vocabulary required to reload the model
- "voice_refs/" β€” reference voice clips for voice cloning
- "app.py" β€” Gradio demo application
- "src/inference.py" β€” reusable inference code
- "requirements.txt" β€” Python dependencies

πŸ“ Repository structure

.
β”œβ”€β”€ app.py
β”œβ”€β”€ config.json
β”œβ”€β”€ model.pth
β”œβ”€β”€ vocab.json
β”œβ”€β”€ requirements.txt
β”œβ”€β”€ src/
β”‚   └── inference.py
└── voice_refs/
    β”œβ”€β”€ slt.wav
    β”œβ”€β”€ bdl.wav
    └── ...

🐍 Usage outside this Space

The reusable inference interface can be used directly from Python:

from src.inference import generate

path = generate(
    text="Hello, this is Flex-Audio speaking.",
    language="en",
    voice="slt",  # filename without .wav from voice_refs/
)

Load the model directly

from TTS.tts.configs.xtts_config import XttsConfig
from TTS.tts.models.xtts import Xtts

config = XttsConfig()
config.load_json("config.json")

model = Xtts.init_from_config(config)
model.load_checkpoint(
    config,
    checkpoint_dir=".",
    eval=True,
)

model.cuda()

outputs = model.synthesize(
    "Your text here",
    config,
    speaker_wav="voice_refs/slt.wav",
    language="en",
)

🌍 Languages

Flex-Audio supports 17 languages with real voice cloning, inherited from XTTS-v2:

en, es, fr, de, it, pt, pl, tr, ru, nl, cs,
ar, zh-cn, ja, hu, ko, hi

🎀 Voices

The "voice_refs/" directory contains the reference voice clips available for cloning.

To use a voice, provide its filename without the ".wav" extension:

voice="slt"

For the complete list of available voices, see the files inside "voice_refs/".

πŸ“œ License

Proprietary / All Rights Reserved.

No permission is granted to use, copy, modify, distribute, sublicense, sell, or deploy the model or its files except as expressly permitted by the copyright holder.

The model also inherits the applicable terms of its base model. Check the "XTTS-v2 license" (https://huggingface.co/coqui/XTTS-v2) before using or redistributing this model.