aufklarer's picture
Document speech-swift Community-1 runtime
a14e6c4 verified
|
Raw
History Blame Contribute Delete
7.2 kB
---
license: cc-by-4.0
language:
- multilingual
tags:
- coreml
- speaker-diarization
- pyannote
- wespeaker
- apple-silicon
base_model: pyannote/speaker-diarization-community-1
library_name: coremltools
pipeline_tag: audio-classification
---
# Pyannote Community-1 Core ML
Precompiled Core ML neural stages and native VBx data for the offline
[Pyannote Community-1](https://huggingface.co/pyannote/speaker-diarization-community-1)
speaker-diarization pipeline on Apple platforms.
> Part of the [soniqo.audio](https://soniqo.audio) speech toolkit. See the
> [speaker-diarization guide](https://soniqo.audio/guides/diarize), inspect the
> [native Swift Community-1 runtime](https://github.com/soniqo/speech-swift/tree/a6ed5a5/Sources/SpeechVAD),
> or browse the [Core ML Speech Models collection](https://huggingface.co/collections/aufklarer/coreml-speech-models-69b52fb6987dffb06fd47570).
This is a pipeline bundle, not one end-to-end model. The host must run powerset
decoding, speaker counting, overlap-aware mask selection, VBx clustering, and
timeline reconstruction exactly as described by `config.json`.
## Model
| Component | Parameters | Precision | Format | Size | Context / sample rate |
|---|---:|---|---|---:|---|
| PyanNet segmentation | 1.49M | FP32 | compiled Core ML | 5.7 MiB | 10 s / 16 kHz mono |
| Masked WeSpeaker | 6.86M | FP32 | compiled Core ML | 26.2 MiB | 10 s + three 589-frame masks |
| PLDA transforms | β€” | FP32 | safetensors | 0.19 MiB | 256 to 128 dimensions |
The WeSpeaker graph includes the exact Kaldi filterbank and weighted statistics
pooling used by Community-1. Its vectors are intentionally not normalized before
PLDA. Both Core ML graphs use fixed batch size one and require iOS 17 or macOS 14
or later.
## Files
| File | Size | Description |
|---|---:|---|
| `segmentation.mlmodelc/` | 5.7 MiB | Precompiled PyanNet segmentation graph |
| `embedding.mlmodelc/` | 26.2 MiB | Precompiled filterbank and masked WeSpeaker graph |
| `plda.safetensors` | 0.19 MiB | Precomputed x-vector and PLDA transforms for VBx |
| `config.json` | β€” | Tensor interfaces, pipeline constants, hashes, and conversion checks |
| `benchmark.json` | β€” | Sanitized aggregate and per-recording benchmark results |
| `LICENSE` | β€” | CC BY 4.0 notice and upstream attribution |
## Performance
Lower diarization error rate (DER) and Jaccard error rate (JER) are better.
Throughput above 1x means faster than realtime.
| Evaluation | DER | JER | Throughput | Interpretation |
|---|---:|---:|---:|---|
| VoxConverse dev, five-file subset, 0.25 s collar, overlap included | 4.65% | 21.42% | 25.0x | Below 10% DER and exactly matches upstream Community-1 |
| speech-swift native runtime, same five files and scorer | 4.66% | 21.43% | 25.2x | 0.02 DER points from the published reference |
| Same five files, strict 0 s collar | 6.99% | 23.07% | 25.0x | Boundary errors are fully counted |
| AMI ES2004a single-meeting diagnostic, known 4 speakers | 23.00% | 23.62% | 24.0x | Exact upstream match; weak absolute accuracy on this meeting |
Tested on Apple M5 Pro with `cpu-and-neural-engine`. Peak process memory
was 863 MiB. The five-file speaker-count estimate
was exact for 3 of 5
recordings, so callers should allow a known or bounded count when available.
The neural stages themselves processed one 10-second window in a median
7.44 ms for segmentation and
33.24 ms for all three masked embeddings.
The benchmark used the official Community-1 host processing and revision
`3533c8cf8e369892e6b79ff1bf80f7b0286a54ee`. Scores and speaker counts matched the upstream PyTorch/MPS
run for every evaluated recording.
The VoxConverse result is a small five-recording release check, not a claim over
the full dataset. The AMI row is one meeting and is shown as a limitation, not a
representative AMI score.
## Swift runtime integration
The matching native runtime is available in
[`soniqo/speech-swift` on `feat/community1-coreml`](https://github.com/soniqo/speech-swift/tree/feat/community1-coreml)
at commit [`a6ed5a5`](https://github.com/soniqo/speech-swift/commit/a6ed5a5).
It runs both Core ML graphs, powerset decoding, speaker counting, PLDA, VBx,
constrained assignment, and timeline reconstruction without Python.
```bash
speech diarize meeting.wav --engine community1
speech diarize meeting.wav --engine community1 --num-speakers 2
speech diarize meeting.wav --engine community1 --min-speakers 2 --max-speakers 6
```
```swift
let diarizer = try await Community1DiarizationPipeline.fromPretrained()
try diarizer.prewarm()
let result = try diarizer.diarize(
audio: samples,
sampleRate: 16_000,
speakerBounds: Community1SpeakerBounds(minimum: 2, maximum: 6)
)
```
The runtime returns diarized segments plus one 256-dimensional centroid for
each detected speaker. Speaker IDs are local to one result; use the centroids
for recording-local or persistent identity matching.
## Download
```bash
hf download aufklarer/Pyannote-Community-1-CoreML --local-dir Pyannote-Community-1-CoreML
```
## Python example
The following runs the segmentation stage. Complete diarization also needs the
host steps and PLDA data described in `config.json`.
```python
import json
from pathlib import Path
import coremltools as ct
import numpy as np
root = Path("Pyannote-Community-1-CoreML")
config = json.loads((root / "config.json").read_text())
model = ct.models.CompiledMLModel(
str(root / config["segmentation"]["model"]),
compute_units=ct.ComputeUnit.CPU_AND_NE,
)
# One 10-second, 16 kHz mono window in [-1, 1].
waveform = np.zeros((1, 1, 160_000), dtype=np.float32)
log_probabilities = model.predict({"waveform": waveform})["log_probabilities"]
print(log_probabilities.shape) # (1, 589, 7)
```
On Apple platforms, load the `.mlmodelc` directories directly. Do not compile an
`.mlpackage` at runtime; compiled artifacts are provided to keep behavior stable
across macOS, iOS, and simulator runtimes.
## Source and license
Derived from
[pyannote/speaker-diarization-community-1](https://huggingface.co/pyannote/speaker-diarization-community-1)
at revision `3533c8cf8e369892e6b79ff1bf80f7b0286a54ee`. Community-1 combines Pyannote segmentation,
WeSpeaker embeddings, and VBx clustering and is distributed under
[CC BY 4.0](https://creativecommons.org/licenses/by/4.0/). Preserve this
attribution when redistributing the bundle.
## Links
- [Speaker diarization guide](https://soniqo.audio/guides/diarize) β€” concepts and public APIs
- [speech-swift](https://github.com/soniqo/speech-swift) β€” Apple speech SDK
- [Native Community-1 runtime](https://github.com/soniqo/speech-swift/tree/a6ed5a5/Sources/SpeechVAD) β€” pinned Swift implementation
- [Community-1 runtime branch](https://github.com/soniqo/speech-swift/tree/feat/community1-coreml) β€” CLI, tests, and documentation
- [Core ML Speech Models](https://huggingface.co/collections/aufklarer/coreml-speech-models-69b52fb6987dffb06fd47570) β€” related Apple bundles
- [Getting started](https://soniqo.audio/getting-started) β€” installation and CLI guide
- [soniqo.audio](https://soniqo.audio) β€” website
- [Blog](https://soniqo.audio/blog) β€” updates and technical articles