hv-sign

Symbolic audio signature. One event stream, six readouts.

The claim in one sentence

hv-sign discovers the symbolic alphabet of a corpus, measures how much of that symbol stream is structure, and reports the distance over which the structure survives.

The unified object

L = (E, A, T, C)

E = event stream          (onset, features, symbol)
A = alphabet              (fixed or discovered)
T = transition structure  (n-gram over A)
C = confidence layers     (meaning gap, decay, alignment)

Everything else is derived. Six readouts, one primitive.

readout what it adds
hv-note fixed alphabet (band, duration, amplitude), 192 symbols
hv-chord polyphonic mode β€” all bands above threshold per event
hv-lex discovered alphabet via k-means on 4-D event features
hv-void structure vs meaning gap over the symbol stream
hv-decay lowest SNR at which the score still recovers
hv-span cross-corpus alignment via aligned centroids

What it produces

Single file (SignReport):

  • events β€” start/end/peak, pitch Hz, band, duration class, amplitude class, fixed symbol, 4-D feature vector
  • fixed_score β€” the sequence of fixed alphabet symbols
  • features β€” the sequence of 4-D feature vectors

Corpus (CorpusSignReport):

  • fixed_counts β€” symbol frequency over the corpus
  • group_counts β€” file counts by filename prefix
  • discovered β€” the learned alphabet: k, centroids, silhouette, full k-search curve
  • void β€” per-file structure / diversity / meaning / gap
  • file_scores β€” symbol sequence for every file

Decay (DecayReport):

  • snr_levels β€” the levels tested, descending
  • noise_rms, onset_floors β€” the mechanism, per level
  • edit_distances β€” normalized Levenshtein between clean and noisy score
  • symbolic_snr β€” the lowest SNR at which the score still recovers
  • note β€” one-line summary

Alignment (SpanReport):

  • centroid_distance β€” mean nearest-centroid distance Aβ†’B
  • histogram_cosine β€” symbol distribution overlap, B projected onto A
  • transition_cosine β€” bigram matrix overlap, B projected onto A
  • alignment β€” weighted combination of the three
  • mapping β€” the Aβ†’B centroid mapping

Install

pip install numpy

Actually β€” no dependencies. Pure stdlib.

## Usage

### Single file

```bash
python hv_sign.py --file ./fsdd/recordings/0_jackson_0.wav
python hv_sign.py --file ./esc50/audio/1-100032-A-0.wav
python hv_sign.py --file x.wav --decay
python hv_sign.py --file x.wav --discover

### Corpus

```bash
python hv_sign.py --dir ./fsdd/recordings --limit 100
python hv_sign.py --dir ./fsdd/recordings --limit 100 --discover
python hv_sign.py --dir ./fsdd/recordings --limit 100 --void

Cross-corpus alignment

python hv_sign.py --align ./fsdd/recordings ./esc50/audio --limit 50

Python

from hv_sign import HVSign

m = HVSign()
report = m.score_file("sample.wav")

print(report.fixed_score)          # [44, 44, 44]
print(report.features)              # [[0.252, 0.511, 0.350, 1.000], ...]

corpus = m.score_directory("./fsdd/recordings", max_files=100)
print(corpus.discovered.k)          # the discovered alphabet size
print(corpus.discovered.silhouette) # how strongly the clusters separate

decay = m.decay_analysis("sample.wav")
print(decay.symbolic_snr)           # lowest SNR at which recovery holds
print(decay.note)                   # 'recovers down to 5 dB'

span = m.align_corpora("./fsdd/recordings", "./esc50/audio", limit=50)
print(span.alignment)               # 0..1

The model

Pipeline

WAV β†’ frames β†’ onsets β†’ events β†’ features β†’ symbols β†’ score
  1. Frame. 20 ms window, 10 ms hop, per-frame RMS.
  2. Onset. Fires when RMS is above onset_floor (0.02) and either rises sharply (Γ— onset_rise_mult, 1.5) or crosses up out of silence. The rising-edge rule prevents a decaying tail from re-triggering; the silence-crossing rule catches slow attacks.
  3. Segment. Extent = connected region above onset_floor. Peak = argmax over the extent only. No fixed search horizon.
  4. Describe. Goertzel power over 16 log-spaced bands (80 Hz – 6 kHz) on a 30 ms window centered on the peak. Top band = argmax. All bands above poly_rel_threshold = the polyphonic band set.
  5. Quantize. Fixed alphabet: band Γ— 12 + dur Γ— 3 + amp, 192 symbols total.
  6. Feature vector. (log10 pitch, log10 duration, peak RMS, log10 rise ratio), each normalized to [0, 1] in a fixed shared space so the discovered alphabet is comparable across corpora.

Fixed alphabet β€” hv-note

axis values
pitch band 16 log-spaced bands, 80 Hz – 6 kHz
duration class <3, 3–7, 8–19, β‰₯20 frames
amplitude class <0.3, 0.3–0.7, β‰₯0.7 of recording peak

Polyphonic mode β€” hv-chord

config.polyphonic = True adds a bands list to each event β€” every band whose Goertzel power is at least poly_rel_threshold (0.25) of the top band. The fixed symbol is unchanged (still top-band). The set-valued alphabet is a v0.2 concern.

Discovered alphabet β€” hv-lex

k-means++ over the 4-D feature vectors, Lloyd iterations until convergence. k chosen by silhouette score in the range [2, min(8, n/3)]. The centroids are the alphabet. k_search reports the full curve so you can see how sharply the optimum is defined.

Structure vs meaning β€” hv-void

Bigram model over the symbol stream. Per file:

  • structure = 1 βˆ’ mean_surprisal / log2(V) β€” how predictable the sequence is
  • diversity = 1 βˆ’ mode_frequency β€” how much the sequence varies
  • meaning = structure Γ— diversity β€” both conditions at once
  • gap = structure βˆ’ meaning β€” how much structure exists without semantics

A repetitive file (all same symbol) has high structure, zero diversity, zero meaning, high gap. A random file has zero structure, zero meaning, zero gap. A patterned file with variation has high structure and high meaning.

Symbolic decay β€” hv-decay

Add Gaussian noise at SNR levels [30, 20, 15, 10, 5, 0]. At each level, raise the effective onset floor to max(config.onset_floor, noise_rms Γ— decay_floor_factor) so the detector doesn't fire on the noise. Re-extract the score. Measure the normalized Levenshtein distance to the clean score.

symbolic_snr is the lowest SNR at which the edit distance stays below decay_edit_threshold (0.3). The note reads recovers down to N dB, or (or lower) if recovery holds at the bottom of the tested range.

Cross-corpus alignment β€” hv-span

Discover an alphabet in each corpus. Map A's centroids to B's by nearest centroid. Project B's symbol histogram and bigram matrix onto A's cluster space via that mapping. Measure:

  • histogram_cosine β€” how similar the symbol distributions are
  • transition_cosine β€” how similar the sequencing patterns are
  • centroid_distance β€” how far the two alphabets sit in feature space
  • alignment = 0.4 Γ— hist + 0.4 Γ— trans + 0.2 Γ— (1 βˆ’ centroid_d)

Benchmarks

Synthesized samples

file events unique fix notes
beeps_low (3Γ— 220 Hz) 3 1 all identical
beeps_high (880/660/440 Hz) 3 3 three distinct bands
chirp_up (200β†’2000 Hz) 1 1 single long event
mix_a (5 varied tones) 5 4 heterogeneous
mix_b (5 varied tones) 5 5 all distinct
hetero (4 varied amp) 4 4 amplitude spread
claps (noise bursts) 4 4 unpitched

Decay, two regimes

Homogeneous file (beeps_low, three identical loud beeps):

SNR noise floor events edit
30 0.007 0.020 3 0.000
20 0.022 0.055 3 0.000
15 0.039 0.097 3 0.000
10 0.069 0.173 3 0.000
5 0.123 0.307 3 0.000
0 0.219 0.546 0 1.000

Heterogeneous file (hetero, four events of mixed amplitude):

SNR noise floor events edit
30 0.005 0.020 4 0.000
20 0.016 0.040 4 0.000
15 0.028 0.071 4 0.250
10 0.050 0.125 3 0.333
5 0.089 0.223 2 0.500
0 0.158 0.396 1 0.750

The homogeneous file loses everything at once. The heterogeneous file loses its quiet events first, then short, then long. Both are correct. The difference is a property of the input.

Cross-corpus alignment

Tonal corpus vs noisy corpus: histogram cos = 0.669, transition cos = 0.243, alignment = 0.498. The corpora share event structure but diverge in sequencing.

When to use it

  • Symbolic search. Query a corpus by symbol pattern. No labels.
  • Corpus discovery. Find the natural alphabet size of a sound collection.
  • Structure vs noise. Separate patterned sequences from unstructured ones without knowing what they mean.
  • Communication range. Estimate the distance at which a call stops being symbolically recoverable.
  • Cross-corpus comparison. Align two sound collections in discovered symbolic space.

When not to use it

  • For speech recognition. No phonemes, no words, no letters.
  • For polyphonic transcription. Polyphonic mode stores band sets but the alphabet and distance metric remain monophonic.
  • For label prediction. The output is symbolic, not semantic.
  • For real-time applications. Pure stdlib k-means and Goertzel are fast on short clips, slow on long passive recordings.
  • For continuous sound without onsets. Drones and hisses produce zero events.

Honest limitations

  • Silhouette is a weak criterion. The k_search curve may have multiple local maxima. When it does, the discovered alphabet size is genuinely ambiguous β€” the data doesn't have a clean natural k.
  • Unpitched events get random bands. White noise has a flat spectrum; the top-band argmax lands wherever the noise happens to be loudest in that 30 ms window. Claps and coin drops produce different symbols on different takes.
  • Amplitude collapses on uniform-loudness files. Amplitude is relative to the recording's own peak. Files whose events are all similar in level produce only amp_class = 2.
  • Decay is step-like on homogeneous files. A file with three identical loud events either recovers all three or recovers none. Intermediate edit distances appear only on files with varied event amplitudes. If you want a smooth decay curve, run the analysis on a file with mixed loudness.
  • Symbolic SNR is capped by the tested range. The default levels stop at 0 dB. Files that still recover at 0 dB are reported as recovers down to 0 dB (or lower) β€” the model has not tested below.
  • transition_cosine requires shared structure. Two corpora with unrelated transition patterns will show transition_cosine β‰ˆ 0 even if they share an alphabet.
  • The alphabet is a k-means result. It is not unique β€” different random seeds can produce different centroids with similar silhouette. The lex_seed config field exists for reproducibility, not for stability.
  • Alignment weights are hand-chosen. 0.4 / 0.4 / 0.2 is plausible but not fit to any human judgment of "how aligned are these two corpora."
  • No calibration against real labeled data. The thresholds are hand-tuned and honest.

Reference

Part of the reader-model series, which has grown beyond reading.

hv-note extracts the symbolic score of a single recording. hv-sign unifies six readouts on the same event stream, discovers the alphabet from the corpus, measures how much of that symbol stream is structure, and reports the range over which the structure survives.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support