You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

1

SigLIP2-ImageShield-90M-224

SigLIP2-ImageShield-90M-224 is a vision-language encoder model fine-tuned from google/siglip2-base-patch16-224 for multi-class image classification. Built on the SiglipForImageClassification architecture, the model is designed to identify and categorize visual content for explicit, suggestive, and safe media filtering.

SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features https://arxiv.org/pdf/2502.14786

Label Space: 5 Classes

The model classifies each image into one of the following content categories:

Class 0: "Anime"
Class 1: "Hentai"
Class 2: "Normal"
Class 3: "Pornography"
Class 4: "Sensual"

Install Dependencies

pip install transformers torch torchvision pillow gradio

Inference Code

import gradio as gr
from transformers import AutoImageProcessor, SiglipForImageClassification
from PIL import Image
import torch

# Load model and processor
model_name = "prithivMLmods/SigLIP2-ImageShield-90M-224"  # Replace with your model path if needed
model = SiglipForImageClassification.from_pretrained(model_name)
processor = AutoImageProcessor.from_pretrained(model_name)

# ID to Label mapping
id2label = {
    "0": "Anime",
    "1": "Hentai",
    "2": "Normal",
    "3": "Pornography",
    "4": "Sensual"
}

def classify_image(image):
    image = Image.fromarray(image).convert("RGB")
    inputs = processor(images=image, return_tensors="pt")

    with torch.no_grad():
        outputs = model(**inputs)
        logits = outputs.logits
        probs = torch.nn.functional.softmax(logits, dim=1).squeeze().tolist()

    prediction = {
        id2label[str(i)]: round(probs[i], 3)
        for i in range(len(probs))
    }

    return prediction

# Gradio Interface
iface = gr.Interface(
    fn=classify_image,
    inputs=gr.Image(type="numpy"),
    outputs=gr.Label(
        num_top_classes=5,
        label="Predicted Content Type"
    ),
    title="SigLIP2-ImageShield-90M",
    description="Classifies images into Anime, Hentai, Normal, Pornography, and Sensual categories."
)

if __name__ == "__main__":
    iface.launch()

Intended Use

This model is intended for applications such as:

  • Content Moderation: Detect explicit or suggestive visual content.
  • Parental Controls: Support AI-based media filtering.
  • Dataset Preprocessing: Categorize and filter image datasets.
  • Online Platforms: Assist with content safety and upload moderation.

Classification Report

Training vs Evaluation Loss / Accuracy

Training vs Evaluation Loss and Accuracy

Precision / Recall / F1-score per Class

Per-Class Precision, Recall, and F1-score

Confusion Matrix

Confusion Matrix

Test Set Class Distribution

Test Set Class Distribution

Overall Prediction Accuracy

Overall Prediction Accuracy

Misalignment Distribution by True Class

Misalignment Distribution by True Class

Acknowledgements

  • Transformers: Transformers provides state-of-the-art machine learning models for text, computer vision, audio, video, and multimodal tasks, supporting both inference and training.

  • SigLIP 2: Multilingual vision-language encoders with improved semantic understanding, localization, and dense feature representations.

Downloads last month
-
Safetensors
Model size
92.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/SigLIP2-ImageShield-90M-224

Finetuned
(124)
this model

Collection including prithivMLmods/SigLIP2-ImageShield-90M-224

Paper for prithivMLmods/SigLIP2-ImageShield-90M-224