| --- |
| license: mit |
| language: |
| - ru |
| tags: |
| - clap |
| - audio-text |
| - speech |
| - russian |
| - contrastive-learning |
| pipeline_tag: feature-extraction |
| --- |
| |
| # Best model: weights + inference script |
|
|
| This package contains the trained weights of the best-performing CLAP model from the paper's |
| multi-seed encoder comparison, plus a standalone script to compute an audio-text similarity score |
| with it — without needing the training pipeline or dataset. |
|
|
| ## Model |
|
|
| - **Text encoder**: `Qwen/Qwen3-Embedding-4B` (frozen, downloaded from HuggingFace). |
| - **Audio encoder**: `openai/whisper-large-v3`, last encoder layer (layer 32/32; frozen). |
| - **Trained parameters** (what's actually in the checkpoint, ~11MB): two projection heads |
| (`mlp3_residual` architecture — Linear → LayerNorm → GELU → Dropout → residual block → Linear, |
| hidden_dim=512, joint_dim=256) mapping the frozen text/audio embeddings into a shared 256-d |
| space, plus a small learned gender-embedding table (male/female/unknown, dim 16, concatenated |
| onto the text embedding before projection) and the InfoNCE logit scale/bias. |
|
|
| ## Checkpoints |
|
|
| All **5 trained seeds** are published (`checkpoints/qwen3_whisper_last_seed{0-4}_best.pt`), matching |
| the paper's mean±std-over-5-seeds methodology — any of them can be used with `infer.py`; they are |
| equivalent for practical use. To reproduce the exact number reported for a specific seed in the |
| paper's tables, use the corresponding `seedN` checkpoint. |
|
|
| The frozen pretrained encoders (Qwen3-Embedding-4B, ~8GB; Whisper-large-v3, ~3GB) are **not** |
| included here — `infer.py` downloads them automatically from HuggingFace on first run. |
|
|
| ## Usage |
|
|
| ```sh |
| pip install -r requirements.txt |
| python infer.py --audio clip.wav --text "тихо прошептала, едва сдерживая слёзы" |
| ``` |
|
|
| Optional flags: |
|
|
| ```sh |
| python infer.py \ |
| --audio clip.wav \ |
| --text "тихо прошептала, едва сдерживая слёзы" \ |
| --gender unknown \ |
| --checkpoint checkpoints/qwen3_whisper_last_seed0_best.pt \ |
| --device cuda |
| ``` |
|
|
| Output: a single number, the cosine similarity between the audio and text joint embeddings, |
| in `[-1, 1]`: |
|
|
| ``` |
| cosine similarity: 0.1823 |
| ``` |
|
|
| `--gender` defaults to `unknown` — this matches the setup used for zero-shot evaluation on |
| external datasets in the paper (reader gender is generally not known/applicable at inference |
| time); pass `male`/`female` explicitly if you want to condition on a specific target voice. |
|
|
| ## Limitations |
|
|
| - `infer.py` is intentionally self-contained: it does not import anything from the main sdclap |
| repository, only reimplements the minimal inference-time logic (see inline comments) — this |
| keeps it runnable with just `pip install -r requirements.txt` and a checkpoint file. |
|
|
| ## License |
|
|
| Code (`infer.py`) — MIT. Checkpoint weights — MIT (trained parameters only, no third-party |
| copyrighted content embedded). |
|
|