File size: 3,368 Bytes
c8d735f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3888423
 
c8d735f
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
---
license: mit
library_name: coreml
pipeline_tag: audio-to-audio
tags:
  - coreml
  - core-ml
  - ios
  - macos
  - apple
  - on-device
  - voice-conversion
  - voice-cloning
  - arxiv:2312.01479
---

# OpenVoice V2 — Core ML

*Voice Cloning*

Zero-shot voice conversion. Clone a speaker from ~10s reference audio.

<p><img src="https://huggingface.co/mlboydaisuke/OpenVoice-V2-CoreML/resolve/main/media/2dfefeb539.gif" alt="OpenVoice V2 demo"></p>

Core ML conversion of [myshell-ai/OpenVoice](https://github.com/myshell-ai/OpenVoice) for on-device inference on iPhone, iPad and Mac. Converted with `coremltools`; the packages are stateless, so all sequencing and buffering lives in your Swift code.

| | |
|---|---|
| Task | audio to audio |
| Upstream | [myshell-ai/OpenVoice](https://github.com/myshell-ai/OpenVoice) |
| Packages | 2 |
| Download size | 58 MB |
| Minimum iOS | 17.0 |
| Peak RAM | ~500 MB |

## Files

| File | Size | Compute units | SHA-256 |
|---|---:|---|---|
| `OpenVoice_SpeakerEncoder.mlpackage.zip` | 1 MB | `cpuAndGPU` | `c3f2a96aaf5ecb5c…` |
| `OpenVoice_VoiceConverter.mlpackage.zip` | 57 MB | `cpuAndGPU` | `ef3ce8a2d1564aef…` |
| **Total** | **58 MB** | | |

`compute_units` is not a suggestion -- it is the configuration the conversion was verified against. Moving a package to a different compute unit can silently change the numerics (FP16 attention overflow) or crash on the GPU.

## Download

```bash
hf download mlboydaisuke/coreml-zoo --include "openvoice/*" --local-dir ./openvoice
unzip './openvoice/openvoice/*.zip' -d ./openvoice
```

## Use in Swift

```swift
import CoreML

let config = MLModelConfiguration()
config.computeUnits = .cpuAndGPU   // as converted — see the table above

// Unzip the .mlpackage, drop it into your Xcode target and Xcode compiles it
// at build time:
let model = try OpenVoice_SpeakerEncoder(configuration: config)

// ...or compile a downloaded .mlpackage at runtime:
let compiled = try await MLModel.compileModel(at: mlpackageURL)
let model = try MLModel(contentsOf: compiled, configuration: config)
```

> This model is split into 2 Core ML packages that are driven in sequence from Swift. Load them one at a time, copy the outputs out of the `MLMultiArray` buffers and release each model before loading the next — two large Core ML models resident at once will OOM on an iPhone.

## Demo

- **Sample app** — [`sample_apps/OpenVoiceDemo`](https://github.com/john-rocky/CoreML-Models/tree/master/sample_apps/OpenVoiceDemo), a standalone SwiftUI project.
- **Models Zoo** — this model is downloadable and runnable inside the [Models Zoo app](https://apps.apple.com/app/id6762083207) on the App Store, no build required.

## Conversion

- Script: [`convert_openvoice.py`](https://github.com/john-rocky/CoreML-Models/blob/master/conversion_scripts/convert_openvoice.py)
- Pitfalls hit during conversion (FP16 overflow, ANE buffer limits, stride handling): [`docs/coreml_conversion_notes.md`](https://github.com/john-rocky/CoreML-Models/blob/master/docs/coreml_conversion_notes.md)
- Model index: [CoreML-Models](https://github.com/john-rocky/CoreML-Models)

## License

The conversion inherits the upstream license: **MIT**.

## Credits

- Upstream authors: [myshell-ai/OpenVoice](https://github.com/myshell-ai/OpenVoice), 2023
- Core ML conversion: john-rocky (Daisuke Majima)