File size: 7,471 Bytes
49bc64a
 
 
 
 
 
 
 
 
 
 
 
14c1b81
 
 
 
 
 
49bc64a
 
 
 
 
 
 
 
df0f54c
 
49bc64a
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
14c1b81
 
 
 
 
 
0dd4e92
14c1b81
0dd4e92
14c1b81
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
0dd4e92
 
 
14c1b81
 
 
0dd4e92
 
 
14c1b81
 
0dd4e92
49bc64a
14c1b81
 
 
49bc64a
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
---
license: cc-by-4.0
tags: [audio-classification, vocal-bursts, non-verbal, voiceclap]
pretty_name: VocalBurst Classifier (single-burst)
---

# VocalBurst Classifier — single-burst

A sharper classifier specialised for clips that contain **exactly one isolated vocal burst** (no speech).
Same taxonomy (82 classes + `no_burst`). **Softmax** (single most-likely type). Use it as the "name the
burst" stage after a detector has isolated a segment.

> ⚠️ **Taxonomy patch (label mapping only, weights unchanged):** the three classes originally trained at
> indices **79 `Hand Scratching Head`**, **80 `Hand Slaps`** and **81 `Slap Face`** are now labelled
> **`No burst`** because of excessive false positives in production. All 83 logits and indices are
> intact. See [Taxonomy patch](#taxonomy-patch-indices-79--80--81-are-now-labelled-no-burst) before
> using raw logits.

## Architecture
`laion/voiceclap-commercial` (768-d, frozen) → **MLP 1024-wide × 3 deep** → 83-way softmax. ~3M params.

## Training
Trained ONLY on **clean single-burst clips**: Gemini-confirmed DACVAE positives (82 classes, exactly one
label, no speech) + pure Emilia `no_burst` clips (~15.9k). Cleaner labels, no speech confusion → sharper.
50 epochs, AdamW, cross-entropy, best-mAP checkpoint.

> **v2 (current):** retrained with **9,000 multilingual burst-free negatives** (FLEURS: Chinese, Hindi, Bengali, Arabic, Persian, Urdu, Tamil, Telugu, Vietnamese, Thai, Indonesian, Japanese, Korean, Swahili, Yoruba, Zulu, Turkish, Russian) — fixes hallucinated bursts on non-European speech.

## Results (held-out pure single-burst clips)
| metric | fine (82) | coarse (16 families) |
|---|---|---|
| macro mAP | **0.68** | **0.87** |
| top-1 exact | 60% | – |
| a true label in top-3 | 88% | – |
On isolated bursts it names the coarse family very reliably (~0.87) and the fine subtype well.

## How it works
Two stages: a **frozen** `laion/voiceclap-commercial` audio encoder turns a clip into a **768-d embedding**
(`encode_waveform`, auto-downloaded — the repo needs no extra setup), then this small trained **MLP head**
maps it to **83 outputs = 82 VocalBurst classes + `no_burst`** (taxonomy: [LAION-AI/voice-taxonomies · vocalburst](https://github.com/LAION-AI/voice-taxonomies/tree/main/vocalburst)).
A **no-burst gate**: if `P(no_burst) ≥ 0.5` the clip is declared burst-free (no false alarm); otherwise the top classes are returned. Clips are truncated to the first 30 s (the encoder's window).

## Usage
```python
from inference import VocalBurstClassifier
clf = VocalBurstClassifier("laion/vocalburst-classifier-single")      # HF repo id, or a local checkout dir
print(clf.predict("clip.wav"))        # -> {no_burst, p_no_burst, top1, predictions:[(class,prob)], group}
```

## Taxonomy patch: indices 79 / 80 / 81 are now labelled `No burst`

> **This section supersedes the earlier "Known caveat: `Slap Face` false positives" note.**
> That note recommended skipping `Slap Face` as top-1 and falling through to the runner-up.
> The problem turned out to be broader than one class and is now fixed in the label mapping
> itself, so the manual skip rule is **no longer needed** — do not apply it on top of this.

### What changed

The model was **originally trained** with three fine classes in the
`hand_and_body_sounds` family:

| index | original class name (as trained) | original group |
|---|---|---|
| **79** | `Hand Scratching Head` | `hand_and_body_sounds` |
| **80** | `Hand Slaps`           | `hand_and_body_sounds` |
| **81** | `Slap Face`            | `hand_and_body_sounds` |

In large-scale production use — re-classifying ~1.5 M CrisperWhisper-detected burst
timestamps from in-the-wild expressive speech — all three fired as top-1 far more often
than is plausible, and manual spot-checks showed the great majority were **false
positives**: the true event was usually a laugh, gasp, sigh, grunt, or simply no burst at
all. `Slap Face` was the worst offender, but `Hand Slaps` and `Hand Scratching Head`
behaved the same way. These three are percussive, broadband, very short events, and the
frozen embedder appears to map a lot of unrelated transient energy (mouth clicks, mic
bumps, plosives, clipping) into that corner of the space.

Accordingly, in `classes.json` **all three labels have been replaced with `"No burst"`**,
and `class_to_group.json` maps `"No burst"` to the group `no_burst`.

### What did *not* change

**The weights were not retrained and not altered.** `model.pt` is bit-identical to the
previous revision. The head still emits **83 logits** and every index keeps its original
position, so **existing checkpoints, cached embeddings, cached logits and stored
`predicted_index` values all stay valid**. This is purely a relabelling of the output
mapping.

### What this means if you use the raw logits

* `argmax` still returns an integer in `[0, 82]`. Indices 79, 80 and 81 now *decode* to
  the string `"No burst"` instead of the three hand-sound names.
* `"No burst"` therefore appears **three times** in `classes.json`. Do not use
  `classes.index("No burst")` to recover an index — it will always give 79. Always keep
  the integer index as the identity and use the list only for display.
* Index **82** is still the model's own dedicated `no_burst` class, spelled
  `no_burst` (lowercase, underscore) to keep the pre-existing gate code working:
  `P(no_burst) ≥ 0.5` still refers to index 82 alone.
* The full set of "this is not a burst" outcomes is therefore
  `{79, 80, 81, 82}`. If you want a single boolean, test membership in that set
  (or test `class_to_group[label] == "no_burst"`), not equality against one index.
* If you need the original three-way hand-sound distinction back (e.g. your domain really
  does contain slaps and you have evaluated it), the information is still in the logits —
  just restore the names at indices 79/80/81 in your own copy of `classes.json`.
* The coarse family `hand_and_body_sounds` now contains only `Finger Snaps` (index 78).

```python
probs = model(emb).softmax(-1)
idx   = int(probs.argmax(-1))          # keep the INDEX as the identity
label = classes[idx]                   # 79/80/81 -> "No burst"
is_burst = idx not in (79, 80, 81, 82)
```

This is an observation from downstream use on in-the-wild speech, not a formal error
analysis on the model's own held-out split — the reported metrics above were computed
**before** this relabelling and still refer to the original 83-way taxonomy.

## Files
`model.pt` (MLP weights, unchanged) · `config.json` (arch) · `classes.json` (83 labels, index order;
79/80/81 relabelled `No burst`) · `class_to_group.json` (fine→coarse families) · `inference.py` ·
`example.py` · `requirements.txt`.

## Interactive demos (audio + predictions)
- Multi-label predictions vs ground truth: https://projects.laion.ai/procedural-voice-captions/vocalburst-predictions/
- Single-burst classifier: https://projects.laion.ai/procedural-voice-captions/vocalburst-single/
- Two-stage detect→classify combo: https://projects.laion.ai/procedural-voice-captions/vocalburst-combo/
- On MOSS character voices: https://projects.laion.ai/procedural-voice-captions/vocalburst-character/

Training data + embeddings: [laion/vocalburst-classification](https://huggingface.co/datasets/laion/vocalburst-classification).
Embedder: [laion/voiceclap-commercial](https://huggingface.co/laion/voiceclap-commercial). License: CC-BY-4.0.