| --- |
| language: |
| - en |
| - zh |
| tags: |
| - onnx |
| - onnxruntime |
| - openvoice |
| - openvoice-v2 |
| - melotts |
| - text-to-speech |
| - voice-conversion |
| - multilingual |
| license: mit |
| pipeline_tag: text-to-speech |
| --- |
| |
| # OpenVoice V2 ONNX (Melo EN/ZH + tone converter) |
|
|
| **Private** ONNX exports for [OpenVoice V2](https://github.com/myshell-ai/OpenVoice) style inference: **MeloTTS** base speech (**English and Chinese** supported via separate Melo exports), **tone-color speaker embedding** from a reference clip, and **tone color conversion** to match the reference voice. Built from a local iOS project under `Docs/OpenVoiceV2/` (`convert_to_onnx.py`, `convert_onnx_precision.py`). |
|
|
| This is **not** an official MyShell release; it is a derivative export for experimentation and on-device ONNX Runtime workflows. |
|
|
| --- |
|
|
| ## Supported languages (this export) |
|
|
| | Layer | English | Chinese (中文) | Notes | |
| |-------|---------|----------------|-------| |
| | **MeloTTS → ONNX** | `melotts_en.onnx` | `melotts_zh.onnx` | Same tensor I/O contract; export both with `OPENVOICE_MELO_LANGS=EN,ZH` (see below). | |
| | **`tone_color_converter.onnx`** | ✓ | ✓ | **Language-agnostic:** linear spectrogram + `sid_src` / `sid_tgt` embeddings. | |
| | **`tone_color_se_extractor.onnx`** | ✓ | ✓ | **Language-agnostic:** reference mel (fixed time in export) → 256-d embedding. | |
| |
| **Base speaker `.pth` (OpenVoiceV2 layout):** e.g. `base_speakers/ses/en-newest.pth` for English-style base timbre, `base_speakers/ses/zh.pth` for Chinese-style base timbre — pick the file that matches your Melo language when doing full OpenVoice-style cloning. |
| |
| **Other Melo languages (ES / FR / JP / KR):** same `convert_to_onnx.py` mechanism with `OPENVOICE_MELO_LANGS`; filenames follow `melotts_{es|fr|jp|kr}.onnx`. |
| |
| **Hub snapshot:** may contain **only** `melotts_en.onnx` until you re-export and re-upload; once `melotts_zh.onnx` is present under `fp32/` (and derived `fp16/` / `int8/`), the same layout applies. |
|
|
| --- |
|
|
| ## Exporting English + Chinese Melo ONNX |
|
|
| From `Docs/OpenVoiceV2` (after checkpoints + deps are installed): |
|
|
| ```bash |
| export OPENVOICE_MELO_LANGS=EN,ZH |
| python convert_to_onnx.py |
| ``` |
|
|
| This produces (alongside shared tone-color ONNX) **`melotts_en.onnx`** and **`melotts_zh.onnx`**. Then run `convert_onnx_precision.py` per directory as usual; precision scripts pick up **every** `melotts_*.onnx` in the input folder. |
|
|
| --- |
|
|
| ## Repository layout |
|
|
| | Path | Precision | Contents | |
| |------|-------------|----------| |
| | `fp32/` | FP32 | `tone_color_se_extractor.onnx`, `tone_color_converter.onnx`, and **one or more** `melotts_*.onnx` (e.g. `melotts_en.onnx`, `melotts_zh.onnx`). | |
| | `fp16/` | Mixed | Same **filenames** as `fp32/`. Global FP16 often breaks ORT on large VITS graphs; the script **falls back to FP32** per file when ORT cannot load the graph. **SE extractor** is often the only true FP16 file. | |
| | `int8/` | Dynamic INT8 (ORT) | Same filenames. **Weight-only dynamic quantization** (`MatMul` / `Gemm`, with `MatMul`-only fallback). | |
|
|
| --- |
|
|
| ## Model I/O (summary) |
|
|
| ### Melo ONNX (`melotts_en.onnx`, `melotts_zh.onnx`, …) |
|
|
| | Input | Shape (typical) | Dtype | |
| |-------|-----------------|-------| |
| | `phone_ids` | `[1, S]` | int64 | |
| | `phone_lens` | `[1]` | int64 | |
| | `sid` | `[1]` | int64 | |
| | `tones` | `[1, S]` | int64 | |
| | `lang_ids` | `[1, S]` | int64 | |
| | `bert_feat` | `[1, 1024, S]` | float32 | |
| | `ja_bert` | `[1, 768, S]` | float32 | |
|
|
| | Output | Shape | Dtype | |
| |--------|-------|-------| |
| | `audio` | `[1, 1, T_wav]` | float32 | |
|
|
| Chinese exports use **Melo’s Chinese text pipeline** (e.g. jieba + multilingual BERT) at **training/export** time; ONNX **inputs remain** the same fields as English. |
|
|
| ### `tone_color_se_extractor.onnx` |
| |
| | Input | Shape | Dtype | |
| |-------|-------|-------| |
| | `mel` | `[1, 513, 128]` (fixed time in export) | float32 | |
| |
| | Output | Shape | Dtype | |
| |--------|-------|-------| |
| | `speaker_embedding` | `[1, 256]` | float32 | |
|
|
| ### `tone_color_converter.onnx` |
|
|
| | Input | Shape | Dtype | |
| |-------|-------|-------| |
| | `y` | `[1, 513, T_mel]` | float32 | |
| | `y_lengths` | `[1]` | int64 | |
| | `sid_src` | `[1, 256, 1]` | float32 | |
| | `sid_tgt` | `[1, 256, 1]` | float32 | |
|
|
| | Output | Shape | Dtype | |
| |--------|-------|-------| |
| | `audio_out` | `[1, 1, T_audio]` | float32 | |
|
|
| Sampling rate and STFT hyperparameters follow **OpenVoice V2 `converter/config.json`** (typically **22050 Hz**, `filter_length=1024`, `hop_length=256`, `win_length=1024`). |
|
|
| --- |
|
|
| ## Intended use |
|
|
| - Research and **on-device** prototyping with **ONNX Runtime** (CPU / optional CoreML EP on Apple platforms). |
| - **English or Chinese** base speech from the matching **`melotts_*.onnx`**, then shared **SE + converter** ONNX. |
| - Pairs with **OpenVoice V2** checkpoints from [myshell-ai/OpenVoiceV2](https://huggingface.co/myshell-ai/OpenVoiceV2) — **not** re-uploaded here. |
| |
| ## Out-of-scope / limitations |
| |
| - **Not** a drop-in replacement for a separate **CoreML** OpenVoice bundle (different tensor contracts in the app). |
| - **No** bundled text frontend: run **Melo-compatible** preprocessing or use `tts_onnx_sentence.py` (`--melo-language EN_NEWEST` or `ZH`, etc.). |
| - **FP16 / INT8:** validate on **your** ORT version and EP. |
|
|
| --- |
|
|
| ## Evaluation |
|
|
| - **Smoke:** `smoke_onnx_openvoice.py`, `smoke_onnx_precision.py` (all `melotts_*.onnx` in the folder). |
| - **Sentence TTS (Python):** `tts_onnx_sentence.py` with `--melo-language` matching the ONNX file. |
|
|
| --- |
|
|
| ## How to download |
|
|
| ```bash |
| pip install huggingface_hub |
| huggingface-cli download aoiandroid/OpenVoiceV2-ONNX-mirror --local-dir ./openvoice_onnx --include "fp32/*" |
| ``` |
|
|
| Use a token with **read** access to this private repo. |
|
|
| --- |
|
|
| ## Citation (upstream) |
|
|
| Cite [MyShell OpenVoice](https://github.com/myshell-ai/OpenVoice) and [MeloTTS](https://github.com/myshell-ai/MeloTTS). This Hub repo is a **convenience export** for on-device experimentation. |
|
|
| --- |
|
|
| ## 日本語サマリー |
|
|
| - **Melo ONNX:** 英語(`melotts_en.onnx`)と中国語(`melotts_zh.onnx`)に対応。`OPENVOICE_MELO_LANGS=EN,ZH` で `convert_to_onnx.py` から両方エクスポート可能(I/O 形式は共通)。 |
| - **コンバータ / SE:** 言語非依存。参照音声で声色を寄せる用途。 |
| - **fp32 / fp16 / int8:** ディレクトリ内の **`melotts_*.onnx` をすべて**精度変換の対象に含めます。 |
| - **公式ではありません。** upstream のライセンス・条件に従ってください。 |
| |