Automatic Speech Recognition
NeMo
English
target-speaker-asr
speaker-diarization
multi-talker
parakeet
rnnt
tdt
Instructions to use BUT-FIT/DiCoP_v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use BUT-FIT/DiCoP_v0.1 with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("BUT-FIT/DiCoP_v0.1") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
File size: 3,407 Bytes
3a73f24 981d979 3a73f24 981d979 3a73f24 c9674b1 3ca6a2a 3a73f24 3ca6a2a | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 | ---
license: cc-by-4.0
language:
- en
library_name: nemo
pipeline_tag: automatic-speech-recognition
base_model: nvidia/parakeet-tdt-0.6b-v2
tags:
- automatic-speech-recognition
- target-speaker-asr
- speaker-diarization
- multi-talker
- parakeet
- nemo
- rnnt
- tdt
---
# DiCoP — Diarization-Conditioned Parakeet
Target-speaker ASR built on [nvidia/parakeet-tdt-0.6b-v2](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2). Given audio and
a diarization, it transcribes **one speaker at a time**.
The conditioning lives inside the encoder. Every frame is labelled silence / target /
non-target / overlap (STNO), and each Conformer layer applies a small learned per-class
transform — an FDDT block — before the layer runs. A whole meeting is decoded per speaker in one
pass: no segmentation, no speaker embeddings, no separation front-end.
| | |
|---|---|
| Base model | [nvidia/parakeet-tdt-0.6b-v2](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2) |
| Parameters | 618M |
| Encoder | 24 × FastConformer, d_model 1024 |
| Encoder frame rate | 12.5 Hz (80 ms) |
| Vocabulary | 1024 BPE tokens |
| Decoder | TDT (token-and-duration transducer) |
| Sample rate | 16000 Hz |
## Usage
This checkpoint **cannot be loaded by `nemo_toolkit` alone**. Its encoder `_target_` points at a
class that lives in the [DiCoP repository](https://github.com/BUTSpeechFIT/DiCoP), and NeMo only resolves `_target_`s
inside the `nemo` package unless that check is relaxed.
```bash
git clone https://github.com/BUTSpeechFIT/DiCoP && cd DiCoP
pip install -r requirements.txt
python infer.py \
--checkpoint BUT-FIT/DiCoP_v0.1 \
--rttm /path/to/rttms/ --audio-dir /path/to/audio/ \
--output hyp.stm
```
To drive the model directly:
```python
import sys
sys.path.insert(0, "/path/to/DiCoP")
from utils.nemo import allow_external_nemo_targets, register_legacy_nemo_aliases
allow_external_nemo_targets()
register_legacy_nemo_aliases()
from src.model.asr_bpe_model import EncDecRNNTBPEModelSTNO
model = EncDecRNNTBPEModelSTNO.from_pretrained("BUT-FIT/DiCoP_v0.1")
```
`transcribe()` is deliberately disabled on this model. NeMo's transcription path cannot supply a
mask, and an unconditioned encoder returns a fluent transcript of whoever is loudest — which
looks correct but is not target-speaker output. Use `infer.py`, or `transcribe_stno()` with an
STNO mask you build yourself (see `src/data/stno.py`).
## Results
Oracle diarization, cpWER and tcpWER (collar 5s) in percent, `whisper_nsf` normalization applied. AMI's half-hour sessions used
windowed local attention
(`-O model.encoder.self_attention_model=rel_pos_local_attn -O model.encoder.att_context_size=[256,256]`)
to bound memory; every other set is full-context, full-session.
| Set | Sessions | cpWER | tcpWER |
|---|---|---|---|
| AMI-SDM dev / test | 18 / 16 | 13.98 / 15.97 | 14.26 / 16.51 |
| AMI-IHM-mix dev / test | 18 / 16 | 11.20 / 11.75 | 11.41 / 12.15 |
| NOTSOFAR-SDM dev1 / eval | 177 / 160 | 17.44 / 17.56 | 17.93 / 17.94 |
| LibriSpeechMix 2mix dev / test | 2703 / 2620 | 2.62 / 2.54 | 2.62 / 2.54 |
| LibriSpeechMix 3mix dev / test | 2703 / 2620 | 6.79 / 6.34 | 6.80 / 6.35 |
| Libri2Mix dev / test clean | 3000 | 4.13 / 4.40 | 4.16 / 4.41 |
| Libri3Mix dev / test clean | 3000 | 30.93 / 33.12 | 31.00 / 33.19 |
## Contact
If you have any questions, reach out to: [iklement@fit.vut.cz](mailto:iklement@fit.vut.cz)
|