You need to agree to share your contact information to access this model
This repository is publicly accessible, but you have to accept the conditions to access its files and content.
Blue-Eye is a content-moderation model. By downloading it you agree to use it under the DINOv3 License and not to use it to monitor or make decisions about individual people.
Log in or Sign Up to review the conditions and access this model content.
Model Card for Blue-Eye
Blue-Eye is an image classifier for content moderation. Given an image, it predicts whether the content is safe, suggestive or explicit, together with a probability for each class. On a benchmark of 3,000 real photographs it reaches 88.9% accuracy, ahead of AWS Rekognition, Gemini 3.1 Pro, Google Cloud Vision and the open-source NSFW detectors it was compared with. It handles both photographs and anime/illustration.
Model Details
Blue-Eye is a DINOv3 ViT-L/16 vision transformer fine-tuned end to end for three-class content classification. The model takes a 512x512 RGB image and returns three class probabilities.
Model Description
- Developed by: Pranshu Patel
- Model type: Vision Transformer image classifier
- Fine-tuned from: facebook/dinov3-vitl16-pretrain-lvd1689m
- Classes:
safe(0),suggestive(1),explicit(2) - Parameters: 303M
- License: DINOv3 License
Model Sources
- Repository: https://github.com/prnshu-p/Blue-Eye
Uses
Direct Use
Blue-Eye is built for moderating sexual content in images:
- filtering explicit or suggestive images out of feeds, search results and timelines
- blurring images or adding content warnings
- age-gating content on platforms that allow adult material
- prioritising images for human moderators
- curating image datasets before training other models
The class definitions follow a nudity and sexual-content rubric:
- safe: no sexualised content, including swimwear, fitness, medical images, breastfeeding and non-sexual art
- suggestive: sexualised but not explicit, such as posed lingerie shots or bare buttocks
- explicit: exposed genitalia, sexual acts or full nudity
Downstream Use
The default prediction is the highest-probability class. Platforms with a stricter or looser policy can
set their own threshold on p(explicit) or p(suggestive) + p(explicit), tuned on their own data. The
model can also be fine-tuned further on a platform's own labels.
Out-of-Scope Use
Blue-Eye covers sexual content only; violence, gore and other policy areas are out of scope. It does not estimate age and is not a tool for detecting child sexual abuse material, for which dedicated hash-matching services should be used. It should not be used to monitor or make decisions about individual people.
Bias, Risks, and Limitations
The boundary between suggestive and its neighbouring classes is the most subjective part of the task, for people and models alike, and it is where most disagreements occur. Performance across demographic groups was not part of this evaluation.
Recommendations
Validate the model on images representative of your own platform before deployment, and keep a human review step for actions that affect user accounts.
How to Get Started with the Model
pip install torch "transformers>=4.56" safetensors pillow numpy huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("pp1618/Blue-Eye")
sys.path.insert(0, path)
from inference import classify
for result in classify(["photo.jpg", "drawing.png"], model=path):
print(result["label"], result["probabilities"])
From the command line:
python inference.py photo.jpg folder_of_images/ --model pp1618/Blue-Eye --device cuda
On GPUs with bfloat16 support, add --precision bf16 (or precision="bf16" in Python) for faster
inference with practically identical predictions.
Training Details
Training Data
About 2.6 million web images covering real photographs and anime/illustration, labelled into the three classes using a commercial content-moderation service, source content ratings and model-assisted relabelling. Evaluation images were removed from the training data.
Training Procedure
Training ran in three progressive fine-tuning stages starting from the DINOv3 ViT-L/16 checkpoint:
- about 1.0M real photographs
- about 1.0M images combining anime/illustration with real photographs
- about 650k class-balanced images
Training regime: 2 epochs per stage at 512x512, AdamW with a one-cycle schedule, label smoothing 0.05, random resized crops, horizontal flips and colour jitter, float32 master weights with bfloat16 mixed precision.
Evaluation
Testing Data and Metrics
The Blue-Eye benchmark contains 3,000 real photographs (1,480 safe, 513 suggestive, 1,007 explicit),
including 525 non-sexual, skin-heavy images such as swimwear, fitness and medical photos. Every system
below was evaluated on the same images with the same three-class labels; commercial services were
queried in August 2026 and their outputs mapped to the three classes. Google Cloud Vision counts an
image as explicit when adult is VERY_LIKELY and as suggestive when racy is VERY_LIKELY; AWS
Rekognition uses its default 50% confidence. The metric is three-class accuracy.
Results
| System | Type | Accuracy |
|---|---|---|
| Blue-Eye | open weights | 88.9% |
| AWS Rekognition | commercial API | 87.3% |
| Gemini 3.1 Pro | commercial model | 86.4% |
| Gemini 3.7 Flash | commercial model | 84.1% |
| Google Cloud Vision SafeSearch | commercial API | 83.5% |
| TostAI/nsfw-image-detection-large | open weights | 79.2% |
| Marqo/nsfw-image-detection-384 | open weights | 73.8% |
| Falconsai/nsfw_image_detection | open weights | 70.9% |
| NudeNet | open weights | 70.4% |
| Freepik/nsfw_image_detector | open weights | 67.6% |
| AdamCodd/vit-base-nsfw-detector | open weights | 62.1% |
Many open-source detectors are binary, so they were also compared on the two binary tasks:
| System | Safe vs not safe | Explicit vs rest |
|---|---|---|
| Blue-Eye | 92.2% | 95.1% |
| Marqo/nsfw-image-detection-384 | 86.4% | 78.3% |
| TostAI/nsfw-image-detection-large | 85.7% | 90.1% |
| Falconsai/nsfw_image_detection | 83.3% | 75.6% |
| Freepik/nsfw_image_detector | 82.6% | 71.8% |
| AdamCodd/vit-base-nsfw-detector | 78.6% | 62.7% |
| NudeNet | 76.9% | 87.5% |
Blue-Eye results across domains:
| Evaluation set | Images | Accuracy |
|---|---|---|
| Real photographs | 4,000 | 91.4% |
| Anime / illustration | 2,000 | 87.2% |
| All | 6,000 | 90.0% |
Per-class recall on the 3,000 real photographs: safe 89.5%, suggestive 75.4%, explicit 94.8%.
Technical Specifications
Model Architecture
- Backbone: DINOv3 ViT-L/16: 24 layers, embedding dimension 1024, 16 heads, 4 register tokens, RoPE
- Pooling: class token concatenated with the mean of the remaining output tokens (2048 features)
- Head: LayerNorm followed by a linear layer to 3 classes
- Input: RGB, shorter edge resized to 537 (bicubic), centre crop to 512x512, ImageNet normalisation
- Weights: float32 safetensors; the head always runs in float32, including under bfloat16 inference
Compute Infrastructure
- Hardware: Google Cloud TPU v5e
- Software: PyTorch, Hugging Face Transformers
License
Blue-Eye is a derivative of DINOv3 and is released under the
DINOv3 License. The full text is in
LICENSE. Commercial use is permitted under its terms.
Citation
BibTeX
@misc{patel2026blueeye,
title = {Blue-Eye: a content-safety image classifier},
author = {Patel, Pranshu},
year = {2026},
url = {https://huggingface.co/pp1618/Blue-Eye}
}
- Downloads last month
- 3
Model tree for pp1618/Blue-Eye
Base model
facebook/dinov3-vit7b16-pretrain-lvd1689m