redimnet2-b6-cn / README.md
OpenASR's picture
docs: align card with sole ReDimNet2 embedder wording
4c8deac verified
|
Raw
History Blame Contribute Delete
4.42 kB
---
license: mit
base_model: PalabraAI/redimnet2
pipeline_tag: feature-extraction
library_name: openasr
tags:
- speaker-diarization
- openasr
- oasr
---
<div align="center">
# ReDimNet2-B6 Speaker Embedder (CN-enhanced) Β· OpenASR
**ReDimNet2-B6 speaker embedder for OpenASR diarization β€” 192-d CN-enhanced embeddings, fully on-device**
[![License](https://img.shields.io/badge/license-MIT-2563eb.svg)](https://github.com/PalabraAI/redimnet2/blob/main/LICENSE)
[![Format](https://img.shields.io/badge/format-.oasr-7c3aed.svg)](https://github.com/QuintinShaw/openasr)
[![Runtime](https://img.shields.io/badge/runtime-OpenASR-111827.svg)](https://openasr.org)
[![Base model](https://img.shields.io/badge/base-redimnet2-f59e0b.svg)](https://huggingface.co/PalabraAI/redimnet2)
Speaker-diarization support pack for the **[OpenASR](https://github.com/QuintinShaw/openasr)** runtime β€”
pure-Rust inference, **no Python at inference time**.
</div>
---
## ✨ Highlights
- πŸ—£οΈ **OpenASR speaker embedder** β€” the only supported speaker-embedding pack for diarization and Voice ID; required for anonymous speaker labels on any ASR family
- 🧬 **192-dim ReDimNet2-B6** β€” PalabraAI's dimension-reshaping speaker net (12.5M params) with a Chinese-enhanced vb2+vox2+cnc2 training mix
- πŸ”’ **Diarization, not identification** β€” anonymous session-relative labels; embeddings stay local and are discarded after the request unless you explicitly enroll a local profile
- 🎯 **Parity-gated packaging** β€” ggml-graph forward pass matches the upstream Python reference at cosine β‰₯ 0.9999 on held-out fixtures
- πŸ¦€ **Native in OpenASR** β€” `.oasr` packs run with no Python at inference, engineered for peak performance on CPU & GPU
## πŸš€ Quickstart
```bash
# 1. Install the OpenASR CLI Β· https://openasr.org
# 2. Pull the pack
openasr pull redimnet2-b6-cn:fp16
# 3. Diarize any transcription (works with every OpenASR ASR model)
openasr transcribe meeting.wav --model xasr-zh-en --diarize --format srt
```
## πŸ“¦ Pack
| Quant | File (`.oasr`) | Size |
|:------|:---------------|-----:|
| fp16 | `redimnet2-b6-cn-fp16.oasr` | 28 MB |
<sub>Single **fp16** build: projection weights ship as fp16; norms/biases and other
parity-sensitive tensors stay f32 inside the pack. No extra public quant tiers.</sub>
## 🧠 About ReDimNet2-B6 Speaker Embedder (CN-enhanced)
ReDimNet2-B6 is PalabraAI's 12.5M-parameter speaker-embedding model from the
ReDimNet2 family, trained on a VoxBlink2 + VoxCeleb2 + CN-Celeb2 mix so English and
Chinese speakers share one embedding space. OpenASR packages the MIT-licensed
checkpoint as a local `.oasr` capability pack and runs it through a ggml graph
(not a pure-Rust hand-written forward). This is the only supported speaker-embedding
stage for diarization and Voice ID: when the pack is missing, diarize/Voice ID
requests fail closed rather than falling back to another embedder. Embeddings are
192-d cosine vectors with a ReDimNet-specific calibration profile.
## βš™οΈ How this pack was made
Converted from [PalabraAI/redimnet2](https://huggingface.co/PalabraAI/redimnet2) with the OpenASR importer:
```bash
openasr model-pack import redimnet2 <src>.safetensors <out>.oasr \
--package-id redimnet2-b6-cn
```
The `.oasr` container is GGUF-backed; projection weights are stored as fp16 while
norms/biases and other parity-sensitive tensors remain f32.
## βš–οΈ License
This pack **inherits the upstream model's license: MIT**
([source](https://github.com/PalabraAI/redimnet2/blob/main/LICENSE)). OpenASR packaging retains the upstream copyright;
the only modification is format conversion.
## πŸ™ Acknowledgements
This pack redistributes **PalabraAI/redimnet2** checkpoint
`b6-vb2+vox2+cnc2_v0-lm.pt` in OpenASR's `.oasr` runtime format. Credit for the
model architecture, training, and original weights belongs to the upstream
PalabraAI / ReDimNet2 authors (paper: ReDimNet2: Scaling Speaker Verification via
Time-Pooled Dimension Reshaping). The upstream model is licensed under **MIT**;
OpenASR packaging retains that license and attribution, with the only modification
being format conversion for local ggml-graph loading.
## πŸ”— Links
- πŸ¦€ **OpenASR** β€” <https://github.com/QuintinShaw/openasr>
- 🌐 **Website** β€” <https://openasr.org>
- πŸ€— **Upstream model** β€” [PalabraAI/redimnet2](https://huggingface.co/PalabraAI/redimnet2)