aufklarer's picture
Publish validated ONNX export
0ce1584 verified
|
Raw
History Blame Contribute Delete
2.3 kB
---
license: mit
language:
- multilingual
tags:
- onnx
- speaker-verification
- speaker-embedding
- voice-identity
base_model: PalabraAI/redimnet2
library_name: onnx
pipeline_tag: audio-classification
---
# ReDimNet2-B6 ONNX Speaker Embeddings
ReDimNet2-B6 produces local speaker embeddings for comparing clean voice
samples. It does not diarize audio or assign names by itself.
## Model
| Property | Value |
|---|---:|
| Parameters | 12.3 million |
| Format | ONNX opset 18, Float32 |
| Model size | 48.9 MiB |
| Input | `audio`, `[1, 96000]` mono Float32 samples |
| Sample rate | 16 kHz |
| Window | 6 seconds |
| Output | `embedding`, `[1, 192]` L2-normalized Float32 |
Applications should repeat clean two-to-six-second speech to fill the input
and center-crop longer samples. Do not use overlapping, mixed, or
unalignable speech as identity evidence.
## Performance
The export is rejected unless its embedding has cosine similarity at least
0.9999 with the pinned PyTorch checkpoint and remains unit-normalized.
| Measurement | Result |
|---|---:|
| Warm six-second CPU inference | 245.5 ms |
Latency is measured on the export host and is not a Windows hardware claim.
## C++ usage
The supported native host is
[speech-core](https://github.com/soniqo/speech-core):
```cpp
#include <speech_core/models/onnx_redimnet_speaker_embedding.h>
speech_core::OnnxReDimNetSpeakerEmbedding model(
"ReDimNet2B6.onnx");
auto embedding = model.embed(samples.data(), samples.size(), 16000);
```
## Files
| File | Description |
|---|---|
| `ReDimNet2B6.onnx` | Fixed-shape speaker encoder |
| `config.json` | Graph contract, provenance, hashes, and parity |
| `README.md` | This model card |
| `LICENSE` | Upstream MIT license |
## Source
Converted from the official
[PalabraAI/ReDimNet2](https://github.com/PalabraAI/redimnet2) B6
`vb2+vox2_v0` large-margin checkpoint. The pinned source revision and
checkpoint SHA-256 are recorded in `config.json`.
Speaker embeddings are useful for labeling; they are not biometric
authentication and do not protect against voice spoofing.
## Links
- [speech-core](https://github.com/soniqo/speech-core) — C++ runtime
- [C++ docs](https://soniqo.audio/speech-core)
- [soniqo.audio](https://soniqo.audio)
- [blog](https://soniqo.audio/blog)