This repository contains Models for an experimental attempt to port diarization to a llama.cpp-based engine. Basic support has been implemented, and model runtimes were created, but development later shifted toward transformer-based diarization models that enable real-time diarization.

The experimental work can be found here:
openresearchtools/engine โ€” experimental-legacy

Older versions, including release 1.5, also include this runtime support.

At the moment, the implementation performs reasonably well for two-speaker diarization, but multi-speaker diarization is not complete. Finishing it would require porting the remaining Python pipeline logic into native C++ or Rust orchestration.

A pre-release build of Transcribe Offline using the same runtime is available here:
Transcribe Offline v1.3.2 pre-release

Please note that results with more than two speakers are currently unusable. This was only an experimental prototype.

This diarization mode is not planned for the full release of the app. The full release will instead use transformer-based diarization.

โš ๏ธ Important Notice

This conversion is an independent, community-driven effort.

It is not affiliated with, endorsed by, or supported by the original authors or the company behind Pyannote.

The original models are licensed under CC-BY-4.0, and this conversion preserves the same license terms.

For details about the original models, please refer to the official Hugging Face repository: https://huggingface.co/pyannote/speaker-diarization-community-1

Downloads last month
9
GGUF
Model size
6.64M params
Architecture
wavtokenizer-dec
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support