YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Marathi + English Emotion Classification — IndicBERT
A Marathi + English emotion classification model fine-tuned from
ai4bharat/IndicBERTv2-MLM-only.
The model predicts one of seven emotions:
- Anger
- Dissatisfaction
- Happiness
- Hope
- Neutral
- Sadness
- Satisfaction
Model Details
- Base Model:
ai4bharat/IndicBERTv2-MLM-only - Task: Emotion Classification
- Languages: Marathi and English
- Number of Classes: 7
- Maximum Sequence Length: 128
- Framework: PyTorch
- Library: Hugging Face Transformers
Emotion Labels
| ID | Emotion |
|---|---|
| 0 | Anger |
| 1 | Dissatisfaction |
| 2 | Happiness |
| 3 | Hope |
| 4 | Neutral |
| 5 | Sadness |
| 6 | Satisfaction |
Training Configuration
| Parameter | Value |
|---|---|
| Epochs | 5 |
| Learning Rate | 2e-5 |
| Train Batch Size | 16 |
| Evaluation Batch Size | 32 |
| Weight Decay | 0.01 |
| Warmup Ratio | 0.10 |
| Gradient Accumulation | 1 |
| Early Stopping Patience | 2 |
| Random Seed | 42 |
| Maximum Sequence Length | 128 |
Evaluation Results
| Metric | Score |
|---|---|
| Accuracy | 0.8354 |
| Macro Precision | 0.8384 |
| Macro Recall | 0.8356 |
| Macro F1 | 0.8350 |
| Weighted Precision | 0.8373 |
| Weighted Recall | 0.8354 |
| Weighted F1 | 0.8343 |
Per-Class Results
| Emotion | Precision | Recall | F1 |
|---|---|---|---|
| Anger | 0.8805 | 0.8468 | 0.8633 |
| Dissatisfaction | 0.8154 | 0.8497 | 0.8322 |
| Happiness | 0.8032 | 0.8706 | 0.8356 |
| Hope | 0.8840 | 0.9056 | 0.8946 |
| Neutral | 0.8578 | 0.6772 | 0.7569 |
| Sadness | 0.8625 | 0.8776 | 0.8700 |
| Satisfaction | 0.7655 | 0.8217 | 0.7926 |
Usage
Install the required packages:
pip install transformers torch
Load the Model
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_id = "omgavali26/emotion"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
print("Model loaded successfully!")
Marathi Emotion Prediction
import torch
text = "मला आज खूप आनंद झाला."
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=128
)
with torch.no_grad():
outputs = model(**inputs)
predicted_id = torch.argmax(outputs.logits, dim=-1).item()
predicted_label = model.config.id2label[predicted_id]
confidence = torch.softmax(
outputs.logits,
dim=-1
)[0][predicted_id].item()
print("Predicted emotion:", predicted_label)
print("Confidence:", confidence)
English Emotion Prediction
import torch
text = "I am very happy today."
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=128
)
with torch.no_grad():
outputs = model(**inputs)
predicted_id = torch.argmax(outputs.logits, dim=-1).item()
predicted_label = model.config.id2label[predicted_id]
confidence = torch.softmax(
outputs.logits,
dim=-1
)[0][predicted_id].item()
print("Predicted emotion:", predicted_label)
print("Confidence:", confidence)
Using the Transformers Pipeline
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="omgavali26/emotion",
tokenizer="omgavali26/emotion"
)
result = classifier("I am very happy today.")
print(result)
Get Scores for All Emotions
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="omgavali26/emotion",
tokenizer="omgavali26/emotion",
top_k=7
)
result = classifier("I am very happy today.")
for item in result[0]:
print(item)
Model Architecture
The model uses:
Input Text
↓
IndicBERT Tokenizer
↓
IndicBERT Encoder
↓
Classification Head
↓
7 Emotion Classes
Supported Emotions
The model predicts:
Anger
Dissatisfaction
Happiness
Hope
Neutral
Sadness
Satisfaction
Long Text
The model was trained with a maximum sequence length of 128 tokens.
For longer text, the original inference workflow uses overlapping chunks with:
- Maximum length: 128
- Overlap: 32 tokens
Long documents should therefore be divided into smaller chunks before prediction.
Intended Use
This model is intended for:
- Marathi emotion classification
- English emotion classification
- Marathi + English text classification
- Emotion analysis
- NLP research
- Academic projects
- Sentiment and emotion related applications
Training Environment
The model was fine-tuned using:
- Python
- PyTorch
- Hugging Face Transformers
- Hugging Face Tokenizers
- Google Colab
Base Model
This model was fine-tuned from:
ai4bharat/IndicBERTv2-MLM-only
Acknowledgements
Thanks to the AI4Bharat team for the IndicBERT model.
Citation
If you use this model in your project or research, please cite the original IndicBERT model and the dataset used for fine-tuning.
- Downloads last month
- 2