Hindi Negative Meme Detection Model πŸ›‘οΈ

A state-of-the-art multimodal AI model for detecting negative/harmful memes in Hindi. This model combines visual understanding (CLIP) with Hindi text understanding (IndicBERT) using attention-based fusion.

Model Description

This model classifies Hindi memes as Negative or Non-Negative based on:

  • Visual content analysis
  • OCR-extracted Hindi/English text
  • Multimodal fusion of both modalities

Architecture

Input: Image + OCR Text
    ↓
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  CLIP ViT-L/14  β”‚   IndicBERT      β”‚
β”‚  (Vision)       β”‚   (Hindi Text)   β”‚
β”‚  1024-dim       β”‚   768-dim        β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
         ↓                 ↓
    Vision Proj       Text Proj
         ↓                 ↓
         β””β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                  ↓
         Attention Fusion (512-dim)
                  ↓
            Classifier Head
                  ↓
      Binary: Negative / Non-Negative

Key Components

  • Vision Encoder: OpenAI CLIP ViT-L/14 (frozen)
  • Text Encoder: AI4Bharat IndicBERT (fine-tuned)
  • Fusion: Multi-head self-attention
  • OCR: EasyOCR (Hindi + English)

Training Details

  • Dataset: 1,141 Hindi memes with labels
  • Epochs: 47
  • Optimizer: AdamW
  • Learning Rate: 3e-5
  • Batch Size: 16
  • Hardware: NVIDIA RTX A6000 (48GB)

Negative Meme Criteria

A meme is labeled as "negative" if it meets ANY of:

  • Sentiment = "Negative"
  • Contains vulgar content
  • Contains abusive language

Usage

Installation

pip install torch torchvision transformers easyocr pillow

Quick Inference

import torch
from PIL import Image
from transformers import AutoModel, AutoTokenizer, CLIPModel
import easyocr

# Load the model
checkpoint = torch.load('best_model.pt', map_location='cuda')

# Initialize components
# ... (see full example in repository)

With the Full Pipeline

from backend.inference import MemePredictor

predictor = MemePredictor('checkpoints/exp2_lr_3e-5/best_model.pt')
result = predictor.predict(image_bytes)

print(f"Is Negative: {result['is_negative']}")
print(f"Confidence: {result['confidence']:.2%}")
print(f"OCR Text: {result['ocr_text']}")

API Usage

Start the FastAPI server:

cd backend
python main.py

Then send requests:

curl -X POST "http://localhost:8000/predict" \
  -H "Content-Type: multipart/form-data" \
  -F "file=@meme.jpg"

Model Files

  • best_model.pt - Full model checkpoint (~4.3GB)
  • config.json - Model configuration
  • src/model.py - Model architecture

Performance

Metric Score
Accuracy ~85%
F1 Score ~0.83

Limitations

  • Optimized for Hindi memes; may not generalize to other languages
  • Requires GPU for efficient inference
  • Large model size (~4.3GB)

Citation

@misc{hindi-meme-detection,
  author = {mrrobot-1001},
  title = {Hindi Negative Meme Detection Model},
  year = {2024},
  publisher = {Hugging Face},
  url = {https://huggingface.co/mrrobot2610/negative-meme-detection-hindi}
}

License

MIT License

Links

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support