Instructions to use autotools/ai_video_studio with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use autotools/ai_video_studio with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf autotools/ai_video_studio:Q4_K_M # Run inference directly in the terminal: llama cli -hf autotools/ai_video_studio:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf autotools/ai_video_studio:Q4_K_M # Run inference directly in the terminal: llama cli -hf autotools/ai_video_studio:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf autotools/ai_video_studio:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf autotools/ai_video_studio:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf autotools/ai_video_studio:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf autotools/ai_video_studio:Q4_K_M
Use Docker
docker model run hf.co/autotools/ai_video_studio:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use autotools/ai_video_studio with Ollama:
ollama run hf.co/autotools/ai_video_studio:Q4_K_M
- Unsloth Studio
How to use autotools/ai_video_studio with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for autotools/ai_video_studio to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for autotools/ai_video_studio to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for autotools/ai_video_studio to start chatting
- Atomic Chat new
- Docker Model Runner
How to use autotools/ai_video_studio with Docker Model Runner:
docker model run hf.co/autotools/ai_video_studio:Q4_K_M
- Lemonade
How to use autotools/ai_video_studio with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull autotools/ai_video_studio:Q4_K_M
Run and chat with the model
lemonade run user.ai_video_studio-Q4_K_M
List all available models
lemonade list
File size: 8,347 Bytes
8dae870 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 | <p align="center">
<img width="1050" height="450" alt="KokoClone Banner" src="https://github.com/user-attachments/assets/26fbb00c-220e-435a-8f54-431781449c76" />
</p>
<h1 align="center">ποΈ KokoClone</h1>
<p align="center">
<a href="https://huggingface.co/spaces/PatnaikAshish/kokoclone">
<img src="https://img.shields.io/badge/π€%20Hugging%20Face-Live%20Demo-blue" alt="Hugging Face Space" />
</a>
<a href="https://huggingface.co/PatnaikAshish/kokoclone">
<img src="https://img.shields.io/badge/π€%20Models-Repository-orange" alt="Hugging Face Models" />
</a>
<img src="https://img.shields.io/badge/Python-3.10%20to%203.12-3776AB.svg?logo=python&logoColor=white" alt="Python" />
<a href="https://opensource.org/licenses/Apache-2.0">
<img src="https://img.shields.io/badge/License-Apache_2.0-green.svg" alt="License" />
</a>
</p>
**KokoClone** is a fast, real-time compatible multilingual voice cloning system built on top of **Kokoro-ONNX**, one of the fastest open-source neural TTS engines available today.
It allows you to:
* **Text β Clone:** Type text in multiple languages, provide a short reference audio clip, and instantly generate speech in that same voice.
* **Audio β Clone:** Re-voice an existing audio recording to sound like any reference speaker β *no transcription needed*.
## Features
### Multilingual Speech Generation
Generate native speech in English (`en`), Hindi (`hi`), French (`fr`), Japanese (`ja`), Chinese (`zh`), Italian (`it`), Portuguese (`pt`), and Spanish (`es`).
### Zero-Shot Voice Cloning
Upload a 3β10 second voice sample and KokoClone instantly transfers its vocal characteristics to the generated speech.
### Audio-to-Audio Voice Conversion
Upload any existing speech recording and re-voice it to sound like a reference speaker. The pipeline skips TTS entirely and runs purely through the Kanade voice-conversion model. Works on recordings of any length thanks to automatic VRAM-aware chunking!
### Automatic Model Handling
On the first run, the required model weights (`.onnx` and `.bin` files) are automatically downloaded from Hugging Face and placed in the correct directories.
### Real-Time Friendly
Built on Kokoro's efficient ONNX runtime pipeline, KokoClone detects your hardware and runs smoothly on both standard laptops (CPU) and workstations (GPU).
## Live Demo
Try it instantly without installing anything:
π **[KokoClone on Hugging Face Spaces](https://huggingface.co/spaces/PatnaikAshish/kokoclone)**
## Installation
You can set up KokoClone using either **Conda** (Recommended) or **uv**.
### 1. Clone the Repository
```bash
git clone https://github.com/Ashish-Patnaik/kokoclone.git
cd kokoclone
```
### 2. Set Up the Environment & Install Dependencies
#### Option A: Using Conda (Recommended)
```bash
conda create -n kokoclone python=3.12.12 -y
conda activate kokoclone
```
**For CPU Users (Mac / Standard Laptops):**
```bash
pip install torch torchaudio --index-url [https://download.pytorch.org/whl/cpu](https://download.pytorch.org/whl/cpu)
pip install -r requirements.txt
```
**For GPU Users (Nvidia GPUs):**
```bash
pip install -r requirements.txt
pip install kokoro-onnx[gpu]
```
#### Option B: Using `uv`
If you prefer [uv](https://docs.astral.sh/uv/) for fast package management:
```bash
# For CPU Users
uv sync
# For GPU Users (Nvidia)
uv sync --extra gpu
# Activate the environment
source .venv/bin/activate # Linux/macOS
.venv\Scripts\activate # Windows
```
## Usage
KokoClone is highly flexible and can be used via Web UI, CLI, or Python API.
### 1. Web Interface (Gradio)
Launch the interactive web app:
```bash
python app.py
```
* **Tab 1 (Text β Clone):** Enter text, pick a language, upload a reference voice, and generate.
* **Tab 2 (Audio β Clone):** Upload source audio and a reference voice, and get back re-voiced audio.
### 2. Command Line Interface (CLI)
Generate speech directly from your terminal.
**Text to cloned speech (default mode):**
```bash
python cli.py --text "Hello from KokoClone" --lang en --ref reference.wav --out output.wav
```
**Audio to re-voiced speech:**
```bash
python cli.py --mode convert --source original_speech.wav --ref target_voice.wav --out revoiced.wav
```
| Argument | Default | Description |
| --- | --- | --- |
| `--mode` | `tts` | `tts` (text β speech) or `convert` (audio β re-voiced audio) |
| `--text` | β | Text to synthesize *(required for `tts` mode)* |
| `--lang` | `en` | Language code: `en hi fr ja zh it es pt` |
| `--source` | β | Path to source audio *(required for `convert` mode)* |
| `--ref` | β | Path to reference voice audio *(always required)* |
| `--out` | `output.wav` | Output file path |
### 3. Python API
Integrate KokoClone into your own Python applications.
**Text to Cloned Speech:**
```python
from core.cloner import KokoClone
cloner = KokoClone()
cloner.generate(
text="This voice is cloned using KokoClone.",
lang="en",
reference_audio="reference.wav",
output_path="output.wav"
)
```
**Audio-to-Audio Voice Conversion:**
```python
import soundfile as sf
from kanade_tokenizer import load_audio
from core.cloner import KokoClone
from core.chunked_convert import chunked_voice_conversion
cloner = KokoClone()
# Load audio tensors
source_wav = load_audio("source_speech.wav", sample_rate=cloner.sample_rate).to(cloner.device)
ref_wav = load_audio("target_voice.wav", sample_rate=cloner.sample_rate).to(cloner.device)
# Convert using VRAM-aware chunking
converted = chunked_voice_conversion(
kanade=cloner.kanade,
vocoder_model=cloner.vocoder,
source_wav=source_wav,
ref_wav=ref_wav,
sample_rate=cloner.sample_rate,
)
sf.write("revoiced_output.wav", converted.numpy(), cloner.sample_rate)
```
## Memory Management for Long Audio
The `chunked_voice_conversion` function in `core/chunked_convert.py` handles memory automatically when converting long audio recordings:
* **VRAM Budget:** On CUDA, chunks are sized so each forward pass uses at most 50% of total GPU memory (configurable via the `vram_fraction` parameter).
* **RoPE Ceiling:** The Kanade `mel_decoder` Transformer has positional embeddings precomputed for 1,024 mel frames. Chunk windows are hard-capped below this limit (β 8.9s of source audio per chunk) with a 10% safety margin to prevent recomputation and quality degradation.
* **Overlap Smoothing:** Each chunk includes a 0.5s overlap on both sides to suppress boundary artifacts.
* **Single-Pass Vocoding:** The full reassembled mel spectrogram is passed to the vocoder in one shot for clean waveform reconstruction.
## Project Structure
```text
app.py β Gradio Web Interface (two-tab UI)
cli.py β Command-line tool (tts and convert modes)
inference.py β Example API usage script
core/
βββ cloner.py β Core TTS + voice cloning engine
βββ chunked_convert.py β VRAM-aware chunked audio conversion
model/ β Downloaded Kokoro model weights (Auto-populates)
voice/ β Downloaded Kokoro voice bins (Auto-populates)
```
## Star History
<a href="https://www.star-history.com/#Ashish-Patnaik/kokoclone&Date">
<picture>
<source media="(prefers-color-scheme: dark)" srcset="https://api.star-history.com/svg?repos=Ashish-Patnaik/kokoclone&type=Date&theme=dark" />
<source media="(prefers-color-scheme: light)" srcset="https://api.star-history.com/svg?repos=Ashish-Patnaik/kokoclone&type=Date" />
<img alt="Star History Chart" src="https://api.star-history.com/svg?repos=Ashish-Patnaik/kokoclone&type=Date" />
</picture>
</a>
## Acknowledgments
This project builds upon the incredible open-source work of:
* **[Kokoro-ONNX](https://github.com/thewh1teagle/kokoro-onnx)** β for fast and efficient neural speech synthesis.
* **[Kanade Tokenizer](https://github.com/frothywater/kanade-tokenizer)** β for the brilliant zero-shot voice conversion architecture.
## License
Licensed under the [Apache 2.0 License](https://www.google.com/search?q=LICENSE).
```
|