Speech-free & music-free source clips — a random sample of 12
Drawn at random (seed 20260809) from the 8,191 clips that survived a two-stage
audio filter over 101,503 source videos from InsAVE-80K and JAVEdit-100k: no speech
(FireRedVAD, then FireRedASR2-AED confirmation on every VAD positive) and no music-related
AudioSet tag anywhere in the top five. These are the inputs to the removal pipeline, not its
outputs. Play with sound on.
101,503 clips screened
31,845 speech-free
8,191 also music-free
12 sampled at random, not curated
Sampled, not cherry-picked — so the failure mode is visible too. 3 of these 12 carry
Speech as their top AudioSet tag even though VAD and ASR both cleared them. That is the
known model disagreement: FireRedASR2 covers Chinese and English, so speech in another language
transcribes empty and slips through. Those clips are flagged in the manifest via
speech_disagreement, and the stricter variant
(kept_no_speech_strict.csv, 6,983 clips) drops them.
insave_general_editing_48600
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Sheep 0.67 Bleat 0.59 Animal 0.57 Goat 0.55 Livestock, farm animals, working animals 0.20
0s voice activity detected
not needed ASR confirmation
— transcript (empty = no speech)
0.0066 audio RMS
insave_eval_00231
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Animal 0.12 Speech 0.11 Door 0.11 Bird 0.07 Chicken, rooster 0.06
0s voice activity detected
not needed ASR confirmation
— transcript (empty = no speech)
0.0161 audio RMS
insave_add_and_remove_04959
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Speech 0.37 Animal 0.07 Sneeze 0.06 Gasp 0.04 Clip-clop 0.03
0.96s voice activity detected
yes ASR confirmation
— transcript (empty = no speech)
0.0069 audio RMS
javedit_7162ff62de8f
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Speech 0.81 Silence 0.08 Television 0.06 Male speech, man speaking 0.05 Inside, small room 0.03
1.23s voice activity detected
yes ASR confirmation
— transcript (empty = no speech)
0.0049 audio RMS
Speech AudioSet says speech
InsAVE-80K
general_editing
1280×704 · 23.98 fps ·
6.715s · mp3 44kHz stereo
insave_general_editing_34794
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Speech 0.66 Animal 0.28 Bird 0.10 Duck 0.05 Quack 0.05
2.18s voice activity detected
yes ASR confirmation
— transcript (empty = no speech)
0.0092 audio RMS
insave_add_and_remove_07472
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Door 0.30 Slam 0.17 Walk, footsteps 0.09 Thunk 0.08 Sliding door 0.05
0.61s voice activity detected
yes ASR confirmation
— transcript (empty = no speech)
0.0206 audio RMS
insave_general_editing_31165
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Liquid 0.15 Rowboat, canoe, kayak 0.11 Boat, Water vehicle 0.09 Stream 0.08 Animal 0.08
0s voice activity detected
not needed ASR confirmation
— transcript (empty = no speech)
0.0029 audio RMS
insave_general_editing_47918
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Frog 0.30 Croak 0.11 Animal 0.05 Walk, footsteps 0.04 Outside, rural or natural 0.04
0s voice activity detected
not needed ASR confirmation
— transcript (empty = no speech)
0.0019 audio RMS
insave_add_and_remove_04160
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Typewriter 0.28 Door 0.27 Sliding door 0.12 Printer 0.04 Cash register 0.04
0s voice activity detected
not needed ASR confirmation
— transcript (empty = no speech)
0.0018 audio RMS
javedit_5b6a0a248356
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Clip-clop 0.63 Animal 0.57 Horse 0.40 Walk, footsteps 0.10 Outside, rural or natural 0.08
0s voice activity detected
not needed ASR confirmation
— transcript (empty = no speech)
0.0024 audio RMS
javedit_4830e7822b49
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Vehicle 0.52 Car 0.13 Truck 0.07 Medium engine (mid frequency) 0.06 Bus 0.05
0s voice activity detected
not needed ASR confirmation
— transcript (empty = no speech)
0.0545 audio RMS
insave_add_and_remove_05404
6 frames, evenly spaced
audio spectrogram (0–8 kHz)
Animal 0.39 Domestic animals, pets 0.27 Bow-wow 0.26 Dog 0.21 Speech 0.15
0.53s voice activity detected
yes ASR confirmation
— transcript (empty = no speech)
0.0332 audio RMS
Manifest: /group2/ct/weihanx/JavisDiT/javisdit/script/extra_train.json — 8,191
entries with dataset attribution, resolvable video paths, AudioSet tags and the per-clip speech
evidence. Clips re-encoded to 360p H.264 / 96 kbps AAC for the web; measurements from the
originals.