Core ML Model Zoo
Collection
PyTorch models converted to Core ML for on-device inference on iPhone, iPad and Mac. β’ 46 items β’ Updated β’ 1
Speaker Identification
Speaker diarization: who spoke when. 16 kHz mono, 10s segments.
Core ML conversion of pyannote/pyannote-audio for on-device inference on iPhone, iPad and Mac. Converted with coremltools; the packages are stateless, so all sequencing and buffering lives in your Swift code.
| Task | voice activity detection |
| Upstream | pyannote/pyannote-audio |
| Packages | 1 |
| Download size | 5 MB |
| Minimum iOS | 17.0 |
| Peak RAM | ~200 MB |
| File | Size | Compute units | SHA-256 |
|---|---|---|---|
SpeakerSegmentation.mlpackage.zip |
5 MB | cpuAndGPU |
dcfa2b98900f2b99β¦ |
| Total | 5 MB |
compute_units is not a suggestion -- it is the configuration the conversion was verified against. Moving a package to a different compute unit can silently change the numerics (FP16 attention overflow) or crash on the GPU.
hf download mlboydaisuke/coreml-zoo --include "diarization/*" --local-dir ./diarization
unzip './diarization/diarization/*.zip' -d ./diarization
import CoreML
let config = MLModelConfiguration()
config.computeUnits = .cpuAndGPU // as converted β see the table above
// Unzip the .mlpackage, drop it into your Xcode target and Xcode compiles it
// at build time:
let model = try SpeakerSegmentation(configuration: config)
// ...or compile a downloaded .mlpackage at runtime:
let compiled = try await MLModel.compileModel(at: mlpackageURL)
let model = try MLModel(contentsOf: compiled, configuration: config)
sample_apps/DiarizationDemo, a standalone SwiftUI project.convert_diarization.pydocs/coreml_conversion_notes.mdThe conversion inherits the upstream license: MIT.
Base model
pyannote/segmentation-3.0