| --- |
| license: other |
| license_name: nvidia-open-model-license |
| license_link: https://developer.nvidia.com/downloads/assets/coreml/nvidia-open-model-license-agreement |
| base_model: nvidia/diar_streaming_sortformer_4spk-v2.1 |
| tags: |
| - onnx |
| - speaker-diarization |
| library_name: onnx |
| --- |
| |
| # Sortformer-Diarization-4spk-ONNX |
|
|
| ONNX export of NVIDIA's streaming Sortformer 4-speaker diarization model, for |
| hosts that drive ONNX Runtime. Apple platforms have a CoreML build; this is the |
| same checkpoint for everything else. |
|
|
| ## Changes from the base model |
|
|
| Exported from `nvidia/diar_streaming_sortformer_4spk-v2.1` with no retraining |
| and no change to weights, architecture or configuration. The export composes |
| the model's pre-encoder and head into one graph per streaming step and traces |
| it at the `default` variant's static shapes. |
|
|
| ## What one call does |
|
|
| ``` |
| in: chunk[1, 3048, 128] mel features, 128 bins |
| chunk_lengths[1] |
| spkcache[1, 188, 512] speaker cache |
| spkcache_lengths[1] |
| fifo[1, 40, 512] |
| fifo_lengths[1] |
| out: spkcache_fifo_chunk_preds[1, 609, 4] per-frame activity, 4 speakers |
| chunk_pre_encode_embs[1, 381, 512] |
| chunk_pre_encode_lengths[1] |
| ``` |
|
|
| One call consumes ~30 seconds of new audio and needs ~3.2 seconds of audio |
| after the chunk it reports on. |
|
|
| ## The host owns the speaker cache |
|
|
| The graph takes `spkcache` and `fifo` as inputs and returns embeddings; it does |
| not update them. Deciding what enters the cache, what waits in the FIFO and |
| what is evicted is the caller's job, and that bookkeeping — the Arrival-Order |
| Speaker Cache — is what makes a speaker index mean the same person for a whole |
| recording. A host that skips it gets per-call speaker numbering and none of the |
| model's advantage over window-local segmenters. |
|
|
| Eviction is not FIFO. Frames are scored by how confidently one speaker and no |
| other is active, boosted twice so every speaker keeps a minimum share and none |
| dominates, then the highest-scoring 188 are kept in chronological order with |
| unused slots filled by a running mean-silence embedding. NeMo's |
| `SortformerModules.streaming_update_async` is the reference. |
|
|
| The cache geometry above belongs to this export. Pairing it with another |
| variant's update period evicts the wrong frames while looking healthy. |
|
|
| ## Measured |
|
|
| Against NeMo's own `forward_streaming` on 132 seconds of two-speaker audio, |
| with the speaker cache overflowing: |
|
|
| | | | |
| |---|---| |
| | mean absolute error | 0.0149 | |
| | decision agreement | 0.9850 | |
| | throughput | 42.4x realtime, 5 calls, 615 ms each | |
|
|
| A bare ONNX Runtime session, without the reference model loaded beside it, is |
| **505 MB resident** and 1251 MB peak, and one call takes **467 ms** to cover |
| 30 s of new audio — **64x realtime** at 12 intra-op threads. The 42.4x above |
| comes from the parity harness, which carries PyTorch and NeMo in the same |
| process; use it to compare against the reference, not to size a deployment. |
|
|
| Measured on a 24-core x86 CPU through ONNX Runtime's CPU provider. The same |
| loop driven by PyTorch instead of this graph gives identical numbers, so the |
| export contributes no error of its own — the difference is between NeMo's two |
| internal streaming paths. |
|
|
| It tracks at most four speakers and degrades beyond that. |
|
|
| ## Files |
|
|
| `sortformer-default.onnx` and `config.json`, which carries the cache sizes and |
| chunk geometry a host needs to size its buffers. |
|
|