Bird-finder β€” offline bird call identifier (v4 + archived v1)

An offline classifier for bird recordings. Point it at a clip and it returns a best guess, a top-3 shortlist, and an honest confidence flag. Nothing leaves your machine: one ONNX file, two PyTorch heads, no network calls.

This repository contains the two openly licensed versions of the model plus the evaluation behind them, so the numbers can be checked rather than taken on trust. Source code: SaiPavankumar22/Bird-Identifier.

Which version do I want?

directory classes params held-out top-1 licence-safe?
v4 β€” publishable (default) models/ 93 (91 species + 2) 1536-d head on a frozen backbone 84.6 % Β± 0.9 yes β€” 0 non-permissive clips
v1 β€” archived starter models_v1_5species/ 5 89,333 43.6 % yes (NPS public domain)

Two intermediate versions (v2, a 95-class CNN, and v3, a 95-class Perch head) were trained with 264 CC BY-NC-ND / BY-NC-SA recordings and are not distributed. Their numbers stay in the evaluation below because they are what v4 is measured against.

Usage

python -m pip install -r requirements.txt      # needs ffmpeg on PATH for .mp3/.m4a

python src/identify.py YOUR.mp3                                  # v4 (93 classes)
python src/identify.py --models-dir models_v1_5species YOUR.mp3  # v1

python src/identify.py --json --no-log YOUR.mp3                  # machine-readable

Each directory is self-contained: models/ holds bundle.json + label_map.json

  • heads + the 413 MB perch_v2_no_dft.onnx backbone; models_v1_5species/ holds the full CNN weights.

Results

Held-out metrics, each on its own protocol

v1 v2 v3 v4
classes 5 95 95 93
chance level 20 % 1.1 % 1.1 % 1.1 %
split 1 recording per species 1 recording per species (+ md5 for DCASE) 3-fold grouped CV 3-fold grouped CV
top-1 43.6 % 47.1 % 84.4 % Β± 0.9 84.6 % Β± 0.9
top-5 100 % 67.2 % 96.9 % 96.8 %
macro-F1 β€” 0.405 0.828 0.798
bird/no-bird AUC β€” 0.846 0.941 Β± 0.004 0.946 Β± 0.002
recordist-disjoint top-1 0.463 Β± 0.379 (LOCO) not run 0.837 Β± 0.016 0.833 Β± 0.020
frozen test (65 clips, 5 species) β€” β€” 1.000 1.000
non-permissive clips in the head no yes (264) yes (264) no

The four top-1 numbers are not on one protocol. v1/v2 are single held-out-recording splits on a CNN trained from scratch; v3/v4 are 3-fold grouped cross-validation over a frozen Perch 2.0 backbone. v3 vs v4 is the only controlled comparison β€” same code, same backbone, same folds, only the data differs.

v4 vs v3 β€” read this before claiming a win

v3 v4
training clips 13,615 16,591 (+21.9 %)
top-1, clip-grouped 0.8442 Β± 0.0088 0.8458 Β± 0.0086 β€” a tie
top-1, recordist-disjoint 0.8369 Β± 0.0165 0.8331 Β± 0.0202
bird/no-bird AUC 0.9415 Β± 0.0042 0.9462 Β± 0.0017 (2.4Γ— tighter)
non-permissive clips 264 0

More data did not buy accuracy β€” +0.0016 against a Β±0.0088 fold spread is a tie. What v4 buys is a clean licence and a stabler bird/no-bird head. The proof that the licence is free: retraining with those 264 non-commercial clips reaches 0.8454 versus the shipped 0.8458 β€” βˆ’0.0004.

Full tables, per-head breakdowns and raw JSON: COMPARISON.md and evaluation/.

Behaviour on audio it should decline

11 clips (2 held-out species, 5 unknown species, 4 synthetic non-bird sounds), from evaluation/compare_all.json:

model non-bird β†’ no bird unknown flagged low held-out correct
v1 0/4 1/5 1/2
v2 1/4 4/5 1/2
v3 4/4 3/5 2/2
v4 4/4 3/5 2/2

The unfavourable half, stated plainly: v3 and v4 still name a species on 5/5 unknown clips. They hedge with the confidence flag instead of refusing. v1 did this while claiming confident every time; v4 flags most of them low.

Training data

source clips licence
iNaturalist 11,605 CC0
DCASE 2018 Task B 6,600 CC BY 4.0
NPS Rocky Mountain sound library 18 public domain (US Gov)
British birdsong dataset 264 CC BY-NC-ND / BY-NC-SA β€” excluded from v4 and v1

The audio itself is not in this repository. It is published separately as SaiPavankumar22/bird-sounds, which contains only the 18,219 permissive files (18.3 GB).

Licence and references

This repository combines works under several licences. Honour all of them.

Code (src/, tools/, eval/) β€” MIT

See LICENSE. Copyright (c) the Bird-finder authors.

Frozen backbone perch_v2_no_dft.onnx β€” Apache-2.0

  • Perch 2.0, Google Research β€” an audio embedding model trained on ~10k species. Redistributed here unmodified.
  • Obtained via justinchuby/Perch-onnx, the ONNX conversion.
  • Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0
  • A copy of the Apache-2.0 terms is in this repository's LICENSE file; the upstream NOTICE/attributions travel with the source project above.
  • Any derived model (including the v3/v4 heads in this repo) must preserve this notice.

Trained heads (species_head.pt, bird_head.pt) β€” MIT

Parameters learned from CC0 / CC BY 4.0 / public-domain audio. No training audio is redistributed in this repository.

Training data β€” per-clip, see the dataset repository

Required attribution for the DCASE portion is reproduced verbatim in the dataset repository's README; reuse of that audio must carry it.

Choosing a licence downstream

if you use you must
the code keep the MIT copyright notice
perch_v2_no_dft.onnx keep the Apache-2.0 notice
v4 or v1 weights nothing beyond the above
the audio in bird-sounds follow each row's license column; DCASE needs credit

Intended use

  • Offline field aid for the 91 species v4 was trained on.
  • A research baseline for "frozen bioacoustic embedding + small head".

Out of scope

  • Not a general bird-identification service. 91 species is a small slice of the birds that exist.
  • Not a detector β€” it assumes there is something to classify.
  • Not a guarantee. A high-confidence flag reports model agreement, not truth.

Files

README.md                ← this card
COMPARISON.md            ← the four-way results write-up
LICENSE                  ← Apache-2.0 (backbone) + MIT (code) notices
requirements.txt
src/                     ← the CLI (identify.py, perch_v3.py)
tools/                   ← shared head definitions + field log
models/                  ← v4 (default): heads + backbone, 93 classes
models_v1_5species/      ← v1: 5-species CNN
evaluation/              ← every raw report behind the numbers above

models/ contains bundle.json, label_map.json, label_index.json, species_head.pt, bird_head.pt, perch_v2_no_dft.onnx.

How v4 runs

  1. audio β†’ mono float32 @ 32 kHz (the backbone's own input; no mel, no dB)
  2. 5.0 s windows, 2.5 s hop, near-silence dropped (RMS < 5e-4), capped at 12
  3. each window β†’ backbone β†’ 1536-d
  4. windows β†’ mean β†’ L2-normalise (identical to pooled_features at training)
  5. pooled vector β†’ species head (93 logits) + bird/no-bird head (2 logits)

Head classes are imported from tools/train_head.py, not re-declared, so the code that saved the weights is the code that loads them.

First call loads the 413 MB backbone (~25 s on CPU). Batch clips in one process.

Limitations

  • Cross-validation is an upper bound. Perch 2.0 was trained on iNaturalist and Xeno-canto, which are inside these folds. The uncontaminated number is the 65-clip frozen test: top-1 1.000, AUC 0.830, 5 species.
  • 13 species have fewer than 10 clips. Per-class recall is noisy for them.
  • Confidence flags measure agreement, not correctness. v4 named a species for an unknown Steller's Jay at 48 % and flagged it moderately.
  • v4 does not beat v3 on species accuracy β€” they are a statistical tie. Pick v4 for its licence, not its score.

Dataset companion

SaiPavankumar22/bird-sounds β€” the 18,219 permissive training clips with per-row licence, attribution and observation URL.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support