File size: 2,993 Bytes
d40d2bd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 | ---
license: mit
library_name: coreml
pipeline_tag: voice-activity-detection
base_model: pyannote/segmentation-3.0
base_model_relation: quantized
tags:
- coreml
- core-ml
- ios
- macos
- apple
- on-device
- speaker-diarization
- vad
- pyannote
- arxiv:2310.13025
---
# Pyannote Diarization — Core ML
*Speaker Identification*
Speaker diarization: who spoke when. 16 kHz mono, 10s segments.
Core ML conversion of [pyannote/pyannote-audio](https://github.com/pyannote/pyannote-audio) for on-device inference on iPhone, iPad and Mac. Converted with `coremltools`; the packages are stateless, so all sequencing and buffering lives in your Swift code.
| | |
|---|---|
| Task | voice activity detection |
| Upstream | [pyannote/pyannote-audio](https://github.com/pyannote/pyannote-audio) |
| Packages | 1 |
| Download size | 5 MB |
| Minimum iOS | 17.0 |
| Peak RAM | ~200 MB |
## Files
| File | Size | Compute units | SHA-256 |
|---|---:|---|---|
| `SpeakerSegmentation.mlpackage.zip` | 5 MB | `cpuAndGPU` | `dcfa2b98900f2b99…` |
| **Total** | **5 MB** | | |
`compute_units` is not a suggestion -- it is the configuration the conversion was verified against. Moving a package to a different compute unit can silently change the numerics (FP16 attention overflow) or crash on the GPU.
## Download
```bash
hf download mlboydaisuke/coreml-zoo --include "diarization/*" --local-dir ./diarization
unzip './diarization/diarization/*.zip' -d ./diarization
```
## Use in Swift
```swift
import CoreML
let config = MLModelConfiguration()
config.computeUnits = .cpuAndGPU // as converted — see the table above
// Unzip the .mlpackage, drop it into your Xcode target and Xcode compiles it
// at build time:
let model = try SpeakerSegmentation(configuration: config)
// ...or compile a downloaded .mlpackage at runtime:
let compiled = try await MLModel.compileModel(at: mlpackageURL)
let model = try MLModel(contentsOf: compiled, configuration: config)
```
## Demo
- **Sample app** — [`sample_apps/DiarizationDemo`](https://github.com/john-rocky/CoreML-Models/tree/master/sample_apps/DiarizationDemo), a standalone SwiftUI project.
- **Models Zoo** — this model is downloadable and runnable inside the [Models Zoo app](https://apps.apple.com/app/id6762083207) on the App Store, no build required.
## Conversion
- Script: [`convert_diarization.py`](https://github.com/john-rocky/CoreML-Models/blob/master/conversion_scripts/convert_diarization.py)
- Pitfalls hit during conversion (FP16 overflow, ANE buffer limits, stride handling): [`docs/coreml_conversion_notes.md`](https://github.com/john-rocky/CoreML-Models/blob/master/docs/coreml_conversion_notes.md)
- Model index: [CoreML-Models](https://github.com/john-rocky/CoreML-Models)
## License
The conversion inherits the upstream license: **MIT**.
## Credits
- Upstream authors: [pyannote/pyannote-audio](https://github.com/pyannote/pyannote-audio), 2021
- Core ML conversion: john-rocky (Daisuke Majima)
|