hv-note

Symbolic audio event extraction. No lexicon. No translation layer.

Given a WAV file, output a score: a sequence of (onset_frame, symbol_id) pairs, where each symbol describes a discrete audio event by (pitch_band, duration_class, amplitude_class).

The claim in one sentence

hv-tempo predicts where the reader slows down. hv-core finds the cheapest edit that breaks an argument. hv-fold counts the passes. hv-note extracts the score β€” the discrete events in a waveform, and nothing else.

What it produces

For a single WAV file:

  • score β€” list of (onset_frame, symbol_id) pairs
  • events β€” for each event: start/end/peak time, pitch Hz, band, duration class, amplitude class, symbol id
  • histogram β€” symbol counts
  • entropy_bits β€” Shannon entropy of the symbol distribution

For a directory:

  • corpus report β€” file count, event count, unique symbols, entropy
  • per-group summary β€” group key extracted from filename
  • file_scores β€” the symbol sequence for every file

Install

pip install numpy

Actually β€” no dependencies. Pure stdlib.

## Usage

### Single file

```bash
python hv_note.py --file ./esc50/audio/1-100032-A-0.wav
python hv_note.py --file ./fsdd/recordings/0_jackson_0.wav

### Directory

```bash
python hv_note.py --dir ./esc50/audio --limit 20
python hv_note.py --dir ./fsdd/recordings

Python

from hv_note import HVNote

m = HVNote()
report = m.score_file("sample.wav")

print(report.score)            # [(0, 44), (24, 44), (49, 44)]
print(report.entropy_bits)     # 0.0
print(report.histogram)        # {44: 3}

fp = m.fingerprint("sample.wav", k=8)
# (44, 44, 44)

CLI flags

python hv_note.py                                  # synthesized demo
python hv_note.py --file x.wav                     # rendered report
python hv_note.py --file x.wav --json              # JSON
python hv_note.py --file x.wav --fingerprint       # just the fingerprint
python hv_note.py --dir ./audio --limit 100        # first 100 files
python hv_note.py --dir ./audio --json             # corpus JSON
python hv_note.py --save-to ./model                # save config

The model

Pipeline

WAV β†’ frames β†’ onsets β†’ events β†’ symbols β†’ score
  1. Frame the signal. 20 ms window, 10 ms hop. Compute per-frame RMS.
  2. Detect onsets. A frame is an onset when its RMS is above onset_floor (0.02) and either rises sharply from the previous frame (Γ— onset_rise_mult) or crosses up out of silence. The rising-edge rule prevents a decaying tail from triggering, and the silence-crossing rule catches slow attacks like chirps.
  3. Segment events. The event's extent is the connected region above onset_floor starting at the onset. The peak is the argmax over that region only β€” never over a fixed search horizon, so a later, louder event cannot hijack the current one.
  4. Describe each event.
    • pitch band β€” argmax of Goertzel power over 16 log-spaced bands (80 Hz – 6 kHz), on a 30 ms window centered on the peak frame.
    • duration class β€” {0: <3 frames, 1: 3–7, 2: 8–19, 3: β‰₯20}.
    • amplitude class β€” {0: <0.3, 1: 0.3–0.7, 2: β‰₯0.7} relative to the recording's own peak RMS.
  5. Encode. symbol = band Β· 12 + dur Β· 3 + amp. 16 Γ— 4 Γ— 3 = 192 symbols.

The alphabet

Fixed, unnamed, universal. Symbol 44 means band 3, duration class 2, amplitude class 2 β€” nothing more. Two recordings producing the same symbol sequence contain the same pattern. No training, no labels, no translation.

Benchmarks

Synthesized samples

signal events unique symbols entropy
3Γ— 220 Hz beeps 3 1 0.000 b
3 mixed beeps (880/660/440 Hz) 3 3 1.585 b
200β†’2000 Hz chirp 1 1 0.000 b
4 noise bursts (claps) 4 4 2.000 b

On real corpora

FSDD (spoken digits, 8 kHz, 3000 recordings). Spoken digits are single long events. Most recordings reduce to 1–3 symbols. The distinguishing signal is pitch band and duration class β€” "one" is short and mid-band; "seven" is longer and higher.

ESC-50 (environmental sound, 44.1 kHz, 2000 recordings). Clips average 8–20 events. The symbol histogram is a signature of the sound category. Confusion appears between categories sharing a dominant pitch band (dog barks and bird calls both light up bands 4–6).

Run either with python hv_note.py --dir <path>. Watch PER-GROUP SUMMARY. Similar unique_sym counts across two groups mean the two categories are symbolically indistinguishable in this 192-symbol alphabet.

When to use it

  • Matching without labels. Compare two recordings by edit distance between their scores. No classifier needed.
  • Repetition detection. Find repeated subsequences within a score.
  • Anomaly detection. Train an n-gram model on a corpus of scores; flag low-likelihood scores.
  • Clustering. Group recordings by symbol histogram.
  • Structural query. "Find all files whose score contains (5,2,1) β†’ (5,2,1) β†’ (3,1,0)."

When not to use it

  • For speech recognition. No phonemes, no words, no letters.
  • For polyphonic music. The model picks a single dominant band per event. Chords collapse.
  • For label prediction. The output is symbolic, not semantic. If you want "this is a dog bark," you need a labeled reference set.
  • For cross-instrument matching. Piano C4 and violin C4 have different timbres, hence different pitch-band energies, hence different symbols.
  • For continuous sound without onsets. Sustained drones, tape hiss, and slow ambient textures produce zero events.

Honest limitations

  • Pitch via Goertzel on 30 ms windows is coarse. Frequency resolution at 80 Hz over 30 ms is ~33 Hz; at 6 kHz it's also ~33 Hz but the log bands are wider, so the top bands are tolerant. Fine for symbolic bucketing, hopeless for tuning.
  • No polyphony. Single argmax per event.
  • Unpitched events get random bands. White noise has a flat spectrum; the pitch-band argmax lands wherever the spectral noise happens to be loudest on that particular 30 ms window. A clap, a snare, and a coin drop all produce different symbols on different takes. If you need to match unpitched sounds, average over multiple takes rather than matching single recordings.
  • Amplitude collapses to a single class on uniform-loudness files. Amplitude is relative to the recording's own peak RMS, so a file whose events are all similar in level produces only amp_class = 2. The amplitude axis is useful within a file that has both loud and quiet events, and carries no information otherwise.
  • Onsets are energy-based. Spectral flux would be sharper, but costs more. This is the cheap version.
  • The alphabet is fixed. A user cannot add custom symbol axes without changing the code. That's the point β€” the alphabet is small and universal.
  • No calibration against annotated corpora. The thresholds are hand-tuned and honest.

Reference

Part of the reader-model series, which has grown beyond reading.

Companion to hv-tempo (pace variation), hv-core (argument robustness), hv-fold (reading passes).

hv-note is the one that doesn't model a reader at all. It models the signal as a symbolic structure and stops there.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support