File size: 7,199 Bytes
d026707 a14e6c4 d026707 a14e6c4 d026707 a14e6c4 d026707 a14e6c4 d026707 a14e6c4 d026707 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 | ---
license: cc-by-4.0
language:
- multilingual
tags:
- coreml
- speaker-diarization
- pyannote
- wespeaker
- apple-silicon
base_model: pyannote/speaker-diarization-community-1
library_name: coremltools
pipeline_tag: audio-classification
---
# Pyannote Community-1 Core ML
Precompiled Core ML neural stages and native VBx data for the offline
[Pyannote Community-1](https://huggingface.co/pyannote/speaker-diarization-community-1)
speaker-diarization pipeline on Apple platforms.
> Part of the [soniqo.audio](https://soniqo.audio) speech toolkit. See the
> [speaker-diarization guide](https://soniqo.audio/guides/diarize), inspect the
> [native Swift Community-1 runtime](https://github.com/soniqo/speech-swift/tree/a6ed5a5/Sources/SpeechVAD),
> or browse the [Core ML Speech Models collection](https://huggingface.co/collections/aufklarer/coreml-speech-models-69b52fb6987dffb06fd47570).
This is a pipeline bundle, not one end-to-end model. The host must run powerset
decoding, speaker counting, overlap-aware mask selection, VBx clustering, and
timeline reconstruction exactly as described by `config.json`.
## Model
| Component | Parameters | Precision | Format | Size | Context / sample rate |
|---|---:|---|---|---:|---|
| PyanNet segmentation | 1.49M | FP32 | compiled Core ML | 5.7 MiB | 10 s / 16 kHz mono |
| Masked WeSpeaker | 6.86M | FP32 | compiled Core ML | 26.2 MiB | 10 s + three 589-frame masks |
| PLDA transforms | β | FP32 | safetensors | 0.19 MiB | 256 to 128 dimensions |
The WeSpeaker graph includes the exact Kaldi filterbank and weighted statistics
pooling used by Community-1. Its vectors are intentionally not normalized before
PLDA. Both Core ML graphs use fixed batch size one and require iOS 17 or macOS 14
or later.
## Files
| File | Size | Description |
|---|---:|---|
| `segmentation.mlmodelc/` | 5.7 MiB | Precompiled PyanNet segmentation graph |
| `embedding.mlmodelc/` | 26.2 MiB | Precompiled filterbank and masked WeSpeaker graph |
| `plda.safetensors` | 0.19 MiB | Precomputed x-vector and PLDA transforms for VBx |
| `config.json` | β | Tensor interfaces, pipeline constants, hashes, and conversion checks |
| `benchmark.json` | β | Sanitized aggregate and per-recording benchmark results |
| `LICENSE` | β | CC BY 4.0 notice and upstream attribution |
## Performance
Lower diarization error rate (DER) and Jaccard error rate (JER) are better.
Throughput above 1x means faster than realtime.
| Evaluation | DER | JER | Throughput | Interpretation |
|---|---:|---:|---:|---|
| VoxConverse dev, five-file subset, 0.25 s collar, overlap included | 4.65% | 21.42% | 25.0x | Below 10% DER and exactly matches upstream Community-1 |
| speech-swift native runtime, same five files and scorer | 4.66% | 21.43% | 25.2x | 0.02 DER points from the published reference |
| Same five files, strict 0 s collar | 6.99% | 23.07% | 25.0x | Boundary errors are fully counted |
| AMI ES2004a single-meeting diagnostic, known 4 speakers | 23.00% | 23.62% | 24.0x | Exact upstream match; weak absolute accuracy on this meeting |
Tested on Apple M5 Pro with `cpu-and-neural-engine`. Peak process memory
was 863 MiB. The five-file speaker-count estimate
was exact for 3 of 5
recordings, so callers should allow a known or bounded count when available.
The neural stages themselves processed one 10-second window in a median
7.44 ms for segmentation and
33.24 ms for all three masked embeddings.
The benchmark used the official Community-1 host processing and revision
`3533c8cf8e369892e6b79ff1bf80f7b0286a54ee`. Scores and speaker counts matched the upstream PyTorch/MPS
run for every evaluated recording.
The VoxConverse result is a small five-recording release check, not a claim over
the full dataset. The AMI row is one meeting and is shown as a limitation, not a
representative AMI score.
## Swift runtime integration
The matching native runtime is available in
[`soniqo/speech-swift` on `feat/community1-coreml`](https://github.com/soniqo/speech-swift/tree/feat/community1-coreml)
at commit [`a6ed5a5`](https://github.com/soniqo/speech-swift/commit/a6ed5a5).
It runs both Core ML graphs, powerset decoding, speaker counting, PLDA, VBx,
constrained assignment, and timeline reconstruction without Python.
```bash
speech diarize meeting.wav --engine community1
speech diarize meeting.wav --engine community1 --num-speakers 2
speech diarize meeting.wav --engine community1 --min-speakers 2 --max-speakers 6
```
```swift
let diarizer = try await Community1DiarizationPipeline.fromPretrained()
try diarizer.prewarm()
let result = try diarizer.diarize(
audio: samples,
sampleRate: 16_000,
speakerBounds: Community1SpeakerBounds(minimum: 2, maximum: 6)
)
```
The runtime returns diarized segments plus one 256-dimensional centroid for
each detected speaker. Speaker IDs are local to one result; use the centroids
for recording-local or persistent identity matching.
## Download
```bash
hf download aufklarer/Pyannote-Community-1-CoreML --local-dir Pyannote-Community-1-CoreML
```
## Python example
The following runs the segmentation stage. Complete diarization also needs the
host steps and PLDA data described in `config.json`.
```python
import json
from pathlib import Path
import coremltools as ct
import numpy as np
root = Path("Pyannote-Community-1-CoreML")
config = json.loads((root / "config.json").read_text())
model = ct.models.CompiledMLModel(
str(root / config["segmentation"]["model"]),
compute_units=ct.ComputeUnit.CPU_AND_NE,
)
# One 10-second, 16 kHz mono window in [-1, 1].
waveform = np.zeros((1, 1, 160_000), dtype=np.float32)
log_probabilities = model.predict({"waveform": waveform})["log_probabilities"]
print(log_probabilities.shape) # (1, 589, 7)
```
On Apple platforms, load the `.mlmodelc` directories directly. Do not compile an
`.mlpackage` at runtime; compiled artifacts are provided to keep behavior stable
across macOS, iOS, and simulator runtimes.
## Source and license
Derived from
[pyannote/speaker-diarization-community-1](https://huggingface.co/pyannote/speaker-diarization-community-1)
at revision `3533c8cf8e369892e6b79ff1bf80f7b0286a54ee`. Community-1 combines Pyannote segmentation,
WeSpeaker embeddings, and VBx clustering and is distributed under
[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/). Preserve this
attribution when redistributing the bundle.
## Links
- [Speaker diarization guide](https://soniqo.audio/guides/diarize) β concepts and public APIs
- [speech-swift](https://github.com/soniqo/speech-swift) β Apple speech SDK
- [Native Community-1 runtime](https://github.com/soniqo/speech-swift/tree/a6ed5a5/Sources/SpeechVAD) β pinned Swift implementation
- [Community-1 runtime branch](https://github.com/soniqo/speech-swift/tree/feat/community1-coreml) β CLI, tests, and documentation
- [Core ML Speech Models](https://huggingface.co/collections/aufklarer/coreml-speech-models-69b52fb6987dffb06fd47570) β related Apple bundles
- [Getting started](https://soniqo.audio/getting-started) β installation and CLI guide
- [soniqo.audio](https://soniqo.audio) β website
- [Blog](https://soniqo.audio/blog) β updates and technical articles
|