Pyannote Diarization β€” Core ML

Speaker Identification

Speaker diarization: who spoke when. 16 kHz mono, 10s segments.

Core ML conversion of pyannote/pyannote-audio for on-device inference on iPhone, iPad and Mac. Converted with coremltools; the packages are stateless, so all sequencing and buffering lives in your Swift code.

Task voice activity detection
Upstream pyannote/pyannote-audio
Packages 1
Download size 5 MB
Minimum iOS 17.0
Peak RAM ~200 MB

Files

File Size Compute units SHA-256
SpeakerSegmentation.mlpackage.zip 5 MB cpuAndGPU dcfa2b98900f2b99…
Total 5 MB

compute_units is not a suggestion -- it is the configuration the conversion was verified against. Moving a package to a different compute unit can silently change the numerics (FP16 attention overflow) or crash on the GPU.

Download

hf download mlboydaisuke/coreml-zoo --include "diarization/*" --local-dir ./diarization
unzip './diarization/diarization/*.zip' -d ./diarization

Use in Swift

import CoreML

let config = MLModelConfiguration()
config.computeUnits = .cpuAndGPU   // as converted β€” see the table above

// Unzip the .mlpackage, drop it into your Xcode target and Xcode compiles it
// at build time:
let model = try SpeakerSegmentation(configuration: config)

// ...or compile a downloaded .mlpackage at runtime:
let compiled = try await MLModel.compileModel(at: mlpackageURL)
let model = try MLModel(contentsOf: compiled, configuration: config)

Demo

Conversion

License

The conversion inherits the upstream license: MIT.

Credits

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for mlboydaisuke/pyannote-segmentation-3.0-CoreML

Quantized
(5)
this model

Collection including mlboydaisuke/pyannote-segmentation-3.0-CoreML