File size: 3,407 Bytes
3a73f24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
981d979
3a73f24
 
 
981d979
3a73f24
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c9674b1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3ca6a2a
3a73f24
3ca6a2a
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
---
license: cc-by-4.0
language:
- en
library_name: nemo
pipeline_tag: automatic-speech-recognition
base_model: nvidia/parakeet-tdt-0.6b-v2
tags:
- automatic-speech-recognition
- target-speaker-asr
- speaker-diarization
- multi-talker
- parakeet
- nemo
- rnnt
- tdt
---

# DiCoP — Diarization-Conditioned Parakeet

Target-speaker ASR built on [nvidia/parakeet-tdt-0.6b-v2](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2). Given audio and
a diarization, it transcribes **one speaker at a time**.

The conditioning lives inside the encoder. Every frame is labelled silence / target /
non-target / overlap (STNO), and each Conformer layer applies a small learned per-class
transform — an FDDT block — before the layer runs. A whole meeting is decoded per speaker in one
pass: no segmentation, no speaker embeddings, no separation front-end.

| | |
|---|---|
| Base model | [nvidia/parakeet-tdt-0.6b-v2](https://huggingface.co/nvidia/parakeet-tdt-0.6b-v2) |
| Parameters | 618M |
| Encoder | 24 × FastConformer, d_model 1024 |
| Encoder frame rate | 12.5 Hz (80 ms) |
| Vocabulary | 1024 BPE tokens |
| Decoder | TDT (token-and-duration transducer) |
| Sample rate | 16000 Hz |

## Usage

This checkpoint **cannot be loaded by `nemo_toolkit` alone**. Its encoder `_target_` points at a
class that lives in the [DiCoP repository](https://github.com/BUTSpeechFIT/DiCoP), and NeMo only resolves `_target_`s
inside the `nemo` package unless that check is relaxed.

```bash
git clone https://github.com/BUTSpeechFIT/DiCoP && cd DiCoP
pip install -r requirements.txt

python infer.py \
    --checkpoint BUT-FIT/DiCoP_v0.1 \
    --rttm /path/to/rttms/ --audio-dir /path/to/audio/ \
    --output hyp.stm
```

To drive the model directly:

```python
import sys
sys.path.insert(0, "/path/to/DiCoP")

from utils.nemo import allow_external_nemo_targets, register_legacy_nemo_aliases

allow_external_nemo_targets()
register_legacy_nemo_aliases()

from src.model.asr_bpe_model import EncDecRNNTBPEModelSTNO

model = EncDecRNNTBPEModelSTNO.from_pretrained("BUT-FIT/DiCoP_v0.1")
```

`transcribe()` is deliberately disabled on this model. NeMo's transcription path cannot supply a
mask, and an unconditioned encoder returns a fluent transcript of whoever is loudest — which
looks correct but is not target-speaker output. Use `infer.py`, or `transcribe_stno()` with an
STNO mask you build yourself (see `src/data/stno.py`).

## Results
Oracle diarization, cpWER and tcpWER (collar 5s) in percent, `whisper_nsf` normalization applied. AMI's half-hour sessions used
windowed local attention
(`-O model.encoder.self_attention_model=rel_pos_local_attn -O model.encoder.att_context_size=[256,256]`)
to bound memory; every other set is full-context, full-session.

| Set | Sessions | cpWER | tcpWER |
|---|---|---|---|
| AMI-SDM dev / test | 18 / 16 | 13.98 / 15.97 | 14.26 / 16.51 |
| AMI-IHM-mix dev / test | 18 / 16 | 11.20 / 11.75 | 11.41 / 12.15 |
| NOTSOFAR-SDM dev1 / eval | 177 / 160 | 17.44 / 17.56 | 17.93 / 17.94 |
| LibriSpeechMix 2mix dev / test | 2703 / 2620 | 2.62 / 2.54 | 2.62 / 2.54 |
| LibriSpeechMix 3mix dev / test | 2703 / 2620 | 6.79 / 6.34 | 6.80 / 6.35 |
| Libri2Mix dev / test clean | 3000 | 4.13 / 4.40 | 4.16 / 4.41 |
| Libri3Mix dev / test clean | 3000 | 30.93 / 33.12 | 31.00 / 33.19 |

## Contact

If you have any questions, reach out to: [iklement@fit.vut.cz](mailto:iklement@fit.vut.cz)