Occlusion Classifiers (DINOv3 + triplet-reduced head)

Binary clear vs occluded classifiers in an embedding + head design: a raw DINOv3 ViT-H+ embedding goes in, a triplet-loss-trained projection reduces it to a 256-d discriminative embedding, and a logistic head reads that. Each self-contained TorchScript .pt records both the reducer and the classifier.

Nine models are published here β€” three projection types x three modalities.

directory projection renders
non_geo/ perspective (non-orthographic) β€”
geo_1k/ orthographic 1k resolution
geo_2k/ orthographic 2k resolution

Each contains occlusion_rgb.pt, occlusion_thermal.pt, and occlusion_matched.pt.

Inputs and outputs

forward(x) takes raw DINOv3 CLS features (N, D) and returns a tuple:

  • embedding (N, 256), L2-normalised β€” the triplet-reduced vector, for re-ID / clustering / viz
  • probs (N, 2) β€” in classes order ["clear", "occluded"]

D depends on the modality (also readable at runtime as m.mu.shape[0]):

modality D features
rgb 1280 RGB CLS
thermal 1280 thermal CLS
matched 2560 [rgb, thermal] concatenated, in that order
import torch
from huggingface_hub import hf_hub_download

path = hf_hub_download("cpraschl/bambi-occlusion-classifiers", "geo_2k/occlusion_matched.pt")
m = torch.jit.load(path)              # no .eval() needed (BatchNorm folded in)

with torch.no_grad():                 # see note below
    emb, probs = m(feats)             # feats: (N, 2560) raw DINOv3 features

label = m.classes[int(probs[0].argmax())]
confidence = float(probs[0].max())

Wrap inference in torch.no_grad(). These models were exported with their parameters still requiring grad, so calling them outside no_grad builds an autograd graph on every call and returns outputs with requires_grad=True β€” you get a warning on float(...) and needless memory growth in a batch loop. no_grad avoids both. (.eval() genuinely is unnecessary β€” BatchNorm is folded in.)

Exported attributes: m.classes (["clear", "occluded"]), m.emb_dim (256), m.reduction ("triplet-loss (batch-hard)").

These models consume DINOv3 features, not images. You must run the DINOv3 ViT-H+ backbone yourself to produce feats. The backbone is separate, unchanged, and not redistributed here β€” see the note on its licensing below.

Architecture

standardize(mu, sd)
  -> Linear(D, 512) + BatchNorm (folded)
  -> ReLU
  -> Linear(512, 256)
  -> L2-normalise            <- triplet loss, batch-hard
  -> logistic head

BatchNorm is fused into the linear layers, so the exported model is deterministic and needs no .eval() call.

Training data

Trained on the BAMBI UAV dataset β€” 389 paired RGB and thermal aerial video sequences recorded by dual-sensor nadir UAVs over Austrian forest habitats, with ~5,100 annotated animal tracks across 12 species classes. The non_geo models use the perspective (as-recorded) frames; geo_1k and geo_2k use orthographic projections of the same material at 1k and 2k render resolution.

Two properties of BAMBI shape the models directly:

  • Tracks are the grouping unit. Evaluation uses GroupKFold by track so that frames from one animal's sequence never straddle the train/test split β€” without this, near-duplicate consecutive frames would inflate accuracy.
  • RGB and thermal are not temporally synchronized in BAMBI. The matched models consume frame-level aligned RGB/thermal pairs, so their (N, 2560) input assumes you have already solved that alignment; they are not a way to avoid it.

Evaluation

Held-out balanced accuracy, GroupKFold by track (grouping by track prevents frames from the same sequence appearing in both train and test):

version rgb thermal matched
non_geo 0.769 0.748 0.845
geo_1k 0.774 0.739 0.812
geo_2k 0.788 0.757 0.804

matched (both modalities fused) is strongest across every projection type. The 2k orthographic renders help rgb (0.774 -> 0.788), mirroring the detectors.

Limitations

  • Feature-space bound. Each model is tied to the exact DINOv3 ViT-H+ features it was trained on. A different backbone, checkpoint, or preprocessing pipeline will silently degrade accuracy β€” the model has no way to detect the mismatch and will still return confident-looking probabilities.
  • Match the model to the projection. The three versions are trained on different projections and are not interchangeable; running perspective features through geo_2k/ (or vice versa) is an out-of-distribution use.
  • Ceiling. Best balanced accuracy is 0.845, so these are useful filters, not authoritative labels. Do not use them where a false "clear" carries real cost without a human in the loop.
  • matched requires both modalities. It expects genuinely paired, co-registered RGB and thermal features; feeding it zeros or an unpaired thermal frame for the missing half is not a supported fallback.

Citation

If you use these models, please also cite the underlying BAMBI dataset β€” the authors ask that work building on it cite their CV4Animals workshop paper (Praschl et al., 2026). See bambi-eco/Dataset for the current citation.

Licensing note on the DINOv3 backbone

These are lightweight heads released here under the MIT license. They are, however, only useful on top of DINOv3, which is distributed by Meta under its own separate license terms that are not MIT and carry their own usage restrictions. The MIT license here covers these heads only β€” it grants no rights to DINOv3. If you intend to use these models, review the DINOv3 license yourself and confirm your use case is permitted under it.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support