File size: 7,199 Bytes
d026707
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a14e6c4
d026707
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a14e6c4
d026707
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a14e6c4
d026707
a14e6c4
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d026707
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a14e6c4
 
d026707
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
---
license: cc-by-4.0
language:
  - multilingual
tags:
  - coreml
  - speaker-diarization
  - pyannote
  - wespeaker
  - apple-silicon
base_model: pyannote/speaker-diarization-community-1
library_name: coremltools
pipeline_tag: audio-classification
---

# Pyannote Community-1 Core ML

Precompiled Core ML neural stages and native VBx data for the offline
[Pyannote Community-1](https://huggingface.co/pyannote/speaker-diarization-community-1)
speaker-diarization pipeline on Apple platforms.

> Part of the [soniqo.audio](https://soniqo.audio) speech toolkit. See the
> [speaker-diarization guide](https://soniqo.audio/guides/diarize), inspect the
> [native Swift Community-1 runtime](https://github.com/soniqo/speech-swift/tree/a6ed5a5/Sources/SpeechVAD),
> or browse the [Core ML Speech Models collection](https://huggingface.co/collections/aufklarer/coreml-speech-models-69b52fb6987dffb06fd47570).

This is a pipeline bundle, not one end-to-end model. The host must run powerset
decoding, speaker counting, overlap-aware mask selection, VBx clustering, and
timeline reconstruction exactly as described by `config.json`.

## Model

| Component | Parameters | Precision | Format | Size | Context / sample rate |
|---|---:|---|---|---:|---|
| PyanNet segmentation | 1.49M | FP32 | compiled Core ML | 5.7 MiB | 10 s / 16 kHz mono |
| Masked WeSpeaker | 6.86M | FP32 | compiled Core ML | 26.2 MiB | 10 s + three 589-frame masks |
| PLDA transforms | β€” | FP32 | safetensors | 0.19 MiB | 256 to 128 dimensions |

The WeSpeaker graph includes the exact Kaldi filterbank and weighted statistics
pooling used by Community-1. Its vectors are intentionally not normalized before
PLDA. Both Core ML graphs use fixed batch size one and require iOS 17 or macOS 14
or later.

## Files

| File | Size | Description |
|---|---:|---|
| `segmentation.mlmodelc/` | 5.7 MiB | Precompiled PyanNet segmentation graph |
| `embedding.mlmodelc/` | 26.2 MiB | Precompiled filterbank and masked WeSpeaker graph |
| `plda.safetensors` | 0.19 MiB | Precomputed x-vector and PLDA transforms for VBx |
| `config.json` | β€” | Tensor interfaces, pipeline constants, hashes, and conversion checks |
| `benchmark.json` | β€” | Sanitized aggregate and per-recording benchmark results |
| `LICENSE` | β€” | CC BY 4.0 notice and upstream attribution |

## Performance

Lower diarization error rate (DER) and Jaccard error rate (JER) are better.
Throughput above 1x means faster than realtime.

| Evaluation | DER | JER | Throughput | Interpretation |
|---|---:|---:|---:|---|
| VoxConverse dev, five-file subset, 0.25 s collar, overlap included | 4.65% | 21.42% | 25.0x | Below 10% DER and exactly matches upstream Community-1 |
| speech-swift native runtime, same five files and scorer | 4.66% | 21.43% | 25.2x | 0.02 DER points from the published reference |
| Same five files, strict 0 s collar | 6.99% | 23.07% | 25.0x | Boundary errors are fully counted |
| AMI ES2004a single-meeting diagnostic, known 4 speakers | 23.00% | 23.62% | 24.0x | Exact upstream match; weak absolute accuracy on this meeting |

Tested on Apple M5 Pro with `cpu-and-neural-engine`. Peak process memory
was 863 MiB. The five-file speaker-count estimate
was exact for 3 of 5
recordings, so callers should allow a known or bounded count when available.

The neural stages themselves processed one 10-second window in a median
7.44 ms for segmentation and
33.24 ms for all three masked embeddings.
The benchmark used the official Community-1 host processing and revision
`3533c8cf8e369892e6b79ff1bf80f7b0286a54ee`. Scores and speaker counts matched the upstream PyTorch/MPS
run for every evaluated recording.

The VoxConverse result is a small five-recording release check, not a claim over
the full dataset. The AMI row is one meeting and is shown as a limitation, not a
representative AMI score.

## Swift runtime integration

The matching native runtime is available in
[`soniqo/speech-swift` on `feat/community1-coreml`](https://github.com/soniqo/speech-swift/tree/feat/community1-coreml)
at commit [`a6ed5a5`](https://github.com/soniqo/speech-swift/commit/a6ed5a5).
It runs both Core ML graphs, powerset decoding, speaker counting, PLDA, VBx,
constrained assignment, and timeline reconstruction without Python.

```bash
speech diarize meeting.wav --engine community1
speech diarize meeting.wav --engine community1 --num-speakers 2
speech diarize meeting.wav --engine community1 --min-speakers 2 --max-speakers 6
```

```swift
let diarizer = try await Community1DiarizationPipeline.fromPretrained()
try diarizer.prewarm()

let result = try diarizer.diarize(
    audio: samples,
    sampleRate: 16_000,
    speakerBounds: Community1SpeakerBounds(minimum: 2, maximum: 6)
)
```

The runtime returns diarized segments plus one 256-dimensional centroid for
each detected speaker. Speaker IDs are local to one result; use the centroids
for recording-local or persistent identity matching.

## Download

```bash
hf download aufklarer/Pyannote-Community-1-CoreML --local-dir Pyannote-Community-1-CoreML
```

## Python example

The following runs the segmentation stage. Complete diarization also needs the
host steps and PLDA data described in `config.json`.

```python
import json
from pathlib import Path

import coremltools as ct
import numpy as np

root = Path("Pyannote-Community-1-CoreML")
config = json.loads((root / "config.json").read_text())
model = ct.models.CompiledMLModel(
    str(root / config["segmentation"]["model"]),
    compute_units=ct.ComputeUnit.CPU_AND_NE,
)

# One 10-second, 16 kHz mono window in [-1, 1].
waveform = np.zeros((1, 1, 160_000), dtype=np.float32)
log_probabilities = model.predict({"waveform": waveform})["log_probabilities"]
print(log_probabilities.shape)  # (1, 589, 7)
```

On Apple platforms, load the `.mlmodelc` directories directly. Do not compile an
`.mlpackage` at runtime; compiled artifacts are provided to keep behavior stable
across macOS, iOS, and simulator runtimes.

## Source and license

Derived from
[pyannote/speaker-diarization-community-1](https://huggingface.co/pyannote/speaker-diarization-community-1)
at revision `3533c8cf8e369892e6b79ff1bf80f7b0286a54ee`. Community-1 combines Pyannote segmentation,
WeSpeaker embeddings, and VBx clustering and is distributed under
[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/). Preserve this
attribution when redistributing the bundle.

## Links

- [Speaker diarization guide](https://soniqo.audio/guides/diarize) β€” concepts and public APIs
- [speech-swift](https://github.com/soniqo/speech-swift) β€” Apple speech SDK
- [Native Community-1 runtime](https://github.com/soniqo/speech-swift/tree/a6ed5a5/Sources/SpeechVAD) β€” pinned Swift implementation
- [Community-1 runtime branch](https://github.com/soniqo/speech-swift/tree/feat/community1-coreml) β€” CLI, tests, and documentation
- [Core ML Speech Models](https://huggingface.co/collections/aufklarer/coreml-speech-models-69b52fb6987dffb06fd47570) β€” related Apple bundles
- [Getting started](https://soniqo.audio/getting-started) β€” installation and CLI guide
- [soniqo.audio](https://soniqo.audio) β€” website
- [Blog](https://soniqo.audio/blog) β€” updates and technical articles