Bird-finder β offline bird call identifier (v4 + archived v1)
An offline classifier for bird recordings. Point it at a clip and it returns a best guess, a top-3 shortlist, and an honest confidence flag. Nothing leaves your machine: one ONNX file, two PyTorch heads, no network calls.
This repository contains the two openly licensed versions of the model plus the evaluation behind them, so the numbers can be checked rather than taken on trust. Source code: SaiPavankumar22/Bird-Identifier.
Which version do I want?
| directory | classes | params | held-out top-1 | licence-safe? | |
|---|---|---|---|---|---|
| v4 β publishable (default) | models/ |
93 (91 species + 2) | 1536-d head on a frozen backbone | 84.6 % Β± 0.9 | yes β 0 non-permissive clips |
| v1 β archived starter | models_v1_5species/ |
5 | 89,333 | 43.6 % | yes (NPS public domain) |
Two intermediate versions (v2, a 95-class CNN, and v3, a 95-class Perch head) were trained with 264 CC BY-NC-ND / BY-NC-SA recordings and are not distributed. Their numbers stay in the evaluation below because they are what v4 is measured against.
Usage
python -m pip install -r requirements.txt # needs ffmpeg on PATH for .mp3/.m4a
python src/identify.py YOUR.mp3 # v4 (93 classes)
python src/identify.py --models-dir models_v1_5species YOUR.mp3 # v1
python src/identify.py --json --no-log YOUR.mp3 # machine-readable
Each directory is self-contained: models/ holds bundle.json + label_map.json
- heads + the 413 MB
perch_v2_no_dft.onnxbackbone;models_v1_5species/holds the full CNN weights.
Results
Held-out metrics, each on its own protocol
| v1 | v2 | v3 | v4 | |
|---|---|---|---|---|
| classes | 5 | 95 | 95 | 93 |
| chance level | 20 % | 1.1 % | 1.1 % | 1.1 % |
| split | 1 recording per species | 1 recording per species (+ md5 for DCASE) | 3-fold grouped CV | 3-fold grouped CV |
| top-1 | 43.6 % | 47.1 % | 84.4 % Β± 0.9 | 84.6 % Β± 0.9 |
| top-5 | 100 % | 67.2 % | 96.9 % | 96.8 % |
| macro-F1 | β | 0.405 | 0.828 | 0.798 |
| bird/no-bird AUC | β | 0.846 | 0.941 Β± 0.004 | 0.946 Β± 0.002 |
| recordist-disjoint top-1 | 0.463 Β± 0.379 (LOCO) | not run | 0.837 Β± 0.016 | 0.833 Β± 0.020 |
| frozen test (65 clips, 5 species) | β | β | 1.000 | 1.000 |
| non-permissive clips in the head | no | yes (264) | yes (264) | no |
The four top-1 numbers are not on one protocol. v1/v2 are single held-out-recording splits on a CNN trained from scratch; v3/v4 are 3-fold grouped cross-validation over a frozen Perch 2.0 backbone. v3 vs v4 is the only controlled comparison β same code, same backbone, same folds, only the data differs.
v4 vs v3 β read this before claiming a win
| v3 | v4 | |
|---|---|---|
| training clips | 13,615 | 16,591 (+21.9 %) |
| top-1, clip-grouped | 0.8442 Β± 0.0088 | 0.8458 Β± 0.0086 β a tie |
| top-1, recordist-disjoint | 0.8369 Β± 0.0165 | 0.8331 Β± 0.0202 |
| bird/no-bird AUC | 0.9415 Β± 0.0042 | 0.9462 Β± 0.0017 (2.4Γ tighter) |
| non-permissive clips | 264 | 0 |
More data did not buy accuracy β +0.0016 against a Β±0.0088 fold spread is a tie. What v4 buys is a clean licence and a stabler bird/no-bird head. The proof that the licence is free: retraining with those 264 non-commercial clips reaches 0.8454 versus the shipped 0.8458 β β0.0004.
Full tables, per-head breakdowns and raw JSON: COMPARISON.md
and evaluation/.
Behaviour on audio it should decline
11 clips (2 held-out species, 5 unknown species, 4 synthetic non-bird sounds),
from evaluation/compare_all.json:
| model | non-bird β no bird |
unknown flagged low | held-out correct |
|---|---|---|---|
| v1 | 0/4 | 1/5 | 1/2 |
| v2 | 1/4 | 4/5 | 1/2 |
| v3 | 4/4 | 3/5 | 2/2 |
| v4 | 4/4 | 3/5 | 2/2 |
The unfavourable half, stated plainly: v3 and v4 still name a species on 5/5 unknown clips. They hedge with the confidence flag instead of refusing. v1 did this while claiming confident every time; v4 flags most of them low.
Training data
| source | clips | licence |
|---|---|---|
| iNaturalist | 11,605 | CC0 |
| DCASE 2018 Task B | 6,600 | CC BY 4.0 |
| NPS Rocky Mountain sound library | 18 | public domain (US Gov) |
| British birdsong dataset | 264 | CC BY-NC-ND / BY-NC-SA β excluded from v4 and v1 |
The audio itself is not in this repository. It is published separately as
SaiPavankumar22/bird-sounds,
which contains only the 18,219 permissive files (18.3 GB).
Licence and references
This repository combines works under several licences. Honour all of them.
Code (src/, tools/, eval/) β MIT
See LICENSE. Copyright (c) the Bird-finder authors.
Frozen backbone perch_v2_no_dft.onnx β Apache-2.0
- Perch 2.0, Google Research β an audio embedding model trained on ~10k species. Redistributed here unmodified.
- Obtained via
justinchuby/Perch-onnx, the ONNX conversion. - Apache License 2.0: https://www.apache.org/licenses/LICENSE-2.0
- A copy of the Apache-2.0 terms is in this repository's
LICENSEfile; the upstream NOTICE/attributions travel with the source project above. - Any derived model (including the v3/v4 heads in this repo) must preserve this notice.
Trained heads (species_head.pt, bird_head.pt) β MIT
Parameters learned from CC0 / CC BY 4.0 / public-domain audio. No training audio is redistributed in this repository.
Training data β per-clip, see the dataset repository
- iNaturalist recordings β CC0. https://www.inaturalist.org Β·
licence recorded per row in the dataset repo's
manifests/inat_clips.csv. - DCASE 2018 Task B / bird-audio-detection β CC BY 4.0. Credit: Mesaros, Heittola, Dikmen, Virtanen, βSound event detection in real environments with application to bird audio detectionβ, DCASE 2018 Workshop. https://dcase.community/challenge2018/task-bird-audio-detection
- NPS Rocky Mountain sound library β public domain (US National Park Service). https://www.nps.gov/subjects/sound/soundlibrary.htm
- Xeno-canto British birdsong (used by the withdrawn v2 and v3 only) β CC BY-NC-ND / BY-NC-SA per recording. https://xeno-canto.org Β· https://www.kaggle.com/datasets/rtatman/british-birdsong-dataset This is why v2 and v3 are not published, and why v4 exists.
Required attribution for the DCASE portion is reproduced verbatim in the dataset repository's README; reuse of that audio must carry it.
Choosing a licence downstream
| if you use | you must |
|---|---|
| the code | keep the MIT copyright notice |
perch_v2_no_dft.onnx |
keep the Apache-2.0 notice |
| v4 or v1 weights | nothing beyond the above |
the audio in bird-sounds |
follow each row's license column; DCASE needs credit |
Intended use
- Offline field aid for the 91 species v4 was trained on.
- A research baseline for "frozen bioacoustic embedding + small head".
Out of scope
- Not a general bird-identification service. 91 species is a small slice of the birds that exist.
- Not a detector β it assumes there is something to classify.
- Not a guarantee. A high-confidence flag reports model agreement, not truth.
Files
README.md β this card
COMPARISON.md β the four-way results write-up
LICENSE β Apache-2.0 (backbone) + MIT (code) notices
requirements.txt
src/ β the CLI (identify.py, perch_v3.py)
tools/ β shared head definitions + field log
models/ β v4 (default): heads + backbone, 93 classes
models_v1_5species/ β v1: 5-species CNN
evaluation/ β every raw report behind the numbers above
models/ contains bundle.json, label_map.json, label_index.json,
species_head.pt, bird_head.pt, perch_v2_no_dft.onnx.
How v4 runs
- audio β mono float32 @ 32 kHz (the backbone's own input; no mel, no dB)
- 5.0 s windows, 2.5 s hop, near-silence dropped (RMS < 5e-4), capped at 12
- each window β backbone β 1536-d
- windows β mean β L2-normalise (identical to
pooled_featuresat training) - pooled vector β species head (93 logits) + bird/no-bird head (2 logits)
Head classes are imported from tools/train_head.py, not re-declared, so the code
that saved the weights is the code that loads them.
First call loads the 413 MB backbone (~25 s on CPU). Batch clips in one process.
Limitations
- Cross-validation is an upper bound. Perch 2.0 was trained on iNaturalist and Xeno-canto, which are inside these folds. The uncontaminated number is the 65-clip frozen test: top-1 1.000, AUC 0.830, 5 species.
- 13 species have fewer than 10 clips. Per-class recall is noisy for them.
- Confidence flags measure agreement, not correctness. v4 named a species for an unknown Steller's Jay at 48 % and flagged it moderately.
- v4 does not beat v3 on species accuracy β they are a statistical tie. Pick v4 for its licence, not its score.
Dataset companion
SaiPavankumar22/bird-sounds
β the 18,219 permissive training clips with per-row licence, attribution and
observation URL.