--- license: other license_name: dinov3-license license_link: https://ai.meta.com/resources/models-and-libraries/dinov3-license base_model: facebook/dinov3-vits16-pretrain-lvd1689m tags: - image-classification - face-detection - pytorch - dinov3 --- # DINOv3 Face-Region Classifier A binary face / non-face classifier finetuned from [`facebook/dinov3-vits16-pretrain-lvd1689m`](https://huggingface.co/facebook/dinov3-vits16-pretrain-lvd1689m) (DINOv3 ViT-S/16) with a single linear head on the pooled embedding. Given a `224x224` crop, it predicts whether that crop contains a face — including non-human, pareidolia-style "faces" (e.g. two dots and a line in tree bark or rock formations) — making it useful as a guidance signal for steering image generation toward face-like content in a specific image region. ## Training data - **ExpWild** and **iCartoonFace**: real human/cartoon face crops - **FacesInThings**: pareidolia face crops (faces seen in inanimate objects/textures) - Negatives: same-size random crops from the same images, sampled to not overlap any labeled face box See [FacePareidolia/classifier](https://github.com/harveymannering/FacePareidolia/tree/main/classifier) for the full training code (`train.py`, `face_dataset.py`). ## Metrics Best checkpoint (epoch 4/20): **val_loss = 0.061, val_acc = 98.0%** (10% image-level holdout). ## Usage The base DINOv3 checkpoint is gated on Hugging Face, so this repo ships its own `config.json` (architecture metadata only — no gated weights) plus the finetuned weights, letting you rebuild the model **without** requesting DINOv3 access: ```python import torch import torch.nn as nn from transformers import AutoConfig, AutoModel from huggingface_hub import hf_hub_download REPO_ID = "harveymannering/DINOv3_FaceRegionClassifier" class Dinov3FaceRegionClassifier(nn.Module): def __init__(self, config): super().__init__() self.backbone = AutoModel.from_config(config) self.head = nn.Linear(config.hidden_size, 1) def forward(self, pixel_values): pooled = self.backbone(pixel_values=pixel_values).pooler_output return self.head(pooled).squeeze(-1) # face logit; sigmoid -> p(face) config = AutoConfig.from_pretrained(REPO_ID) model = Dinov3FaceRegionClassifier(config) weights_path = hf_hub_download(repo_id=REPO_ID, filename="dinov3_face_region_classifier.pth") model.load_state_dict(torch.load(weights_path, map_location="cpu")) model.eval() ``` Input should be `224x224` RGB, normalized with ImageNet mean/std (`mean=[0.485, 0.456, 0.406]`, `std=[0.229, 0.224, 0.225]`). ## License Derived from `facebook/dinov3-vits16-pretrain-lvd1689m`; usage is subject to the terms of that model's license — see the [base model page](https://huggingface.co/facebook/dinov3-vits16-pretrain-lvd1689m) for details.