Learner-sai's picture
Update README.md
4705220 verified
|
Raw
History Blame Contribute Delete
3.5 kB
---
library_name: transformers
license: apache-2.0
base_model: google/muril-base-cased
tags:
- ner
- token-classification
- indic
- multilingual
- marathi
- bengali
- telugu
metrics:
- precision
- recall
- f1
- accuracy
model-index:
- name: muril-ner-multilingual
results: []
---
# muril-ner-multilingual 🏷️
This model is a fine-tuned version of **[google/muril-base-cased](https://huggingface.co/google/muril-base-cased)** for **Named Entity Recognition (NER)** across three major Indic languages: **Marathi (mr)**, **Bengali (bn)**, and **Telugu (te)**.
It uses a joint multilingual full fine-tuning strategy to extract three primary entity types:
- 👤 **PER** (Person)
- 🏢 **ORG** (Organization)
- 📍 **LOC** (Location)
## Model Description
- **Developed by:** Learner-sai
- **Model Type:** Token Classification (NER)
- **Base Model:** `google/muril-base-cased`
- **Languages:** Marathi (`mr`), Bengali (`bn`), Telugu (`te`)
- **License:** Apache 2.0
MuRIL (Multilingual Representations for Indian Languages) was adapted by adding a 7-class linear sequence classification head (`O`, `B-PER`, `I-PER`, `B-ORG`, `I-ORG`, `B-LOC`, `I-LOC`). All parameters were updated during joint training across all three target languages to leverage cross-lingual transfer.
## Intended Uses & Limitations
### Intended Uses 🎯
- Automated entity extraction from news articles, social media, and documents in Marathi, Bengali, and Telugu.
- Downstream NLP applications such as information retrieval, question answering, and knowledge graph construction for Indic languages.
### Limitations ⚠️
- **Grammatical Suffixes:** In agglutinative or highly inflected languages like Marathi and Telugu, locative or case suffixes (e.g., Marathi `-त` in "दिल्लीत") may sometimes be included inside the predicted entity span.
- **Entity Scope:** The model is trained exclusively on `PER`, `ORG`, and `LOC` tags; it will not recognize other categories like dates, monetary values, or product names.
## Training and Evaluation Data
The model was trained on a combined dataset comprising annotated sentences across Marathi, Bengali, and Telugu.
- **Label Schema:** IOB2 format (`B-`, `I-`, `O`) with 7 total classes.
- **Token Alignment:** Subword tokenization artifacts were handled using `-100` label masking on non-initial subwords to ensure clean cross-entropy loss calculation.
## Training Procedure
### Training Hyperparameters
- **Learning Rate:** `3e-05`
- **Train Batch Size:** 16
- **Eval Batch Size:** 32
- **Epochs:** 3
- **Optimizer:** `AdamW (fused)` with $\beta_1=0.9, \beta_2=0.999, \epsilon=1\text{e-}08$
- **LR Scheduler:** Linear with 937 warmup steps (~10% of total steps)
- **Mixed Precision:** Native FP16 (`fp16=True`)
- **Seed:** 42
### Evaluation Metrics
Evaluated on the validation split using `seqeval` (entity-level span matching):
| Epoch | Training Loss | Validation Loss | Precision | Recall | Entity F1 🏆 | Token Accuracy |
| :---: | :-----------: | :-------------: | :-------: | :----: | :---------: | :------------: |
| 1.0 | 0.2845 | 0.2822 | 0.7172 | 0.7725 | 0.7438 | 0.9320 |
| 2.0 | 0.2024 | 0.2335 | 0.7392 | 0.7723 | 0.7554 | 0.9346 |
| **3.0**| **0.1804** | **0.2328** | **0.7382**| **0.7771** | **0.7572** | **0.9352** |
## Framework Versions
- **Transformers:** 5.14.1
- **PyTorch:** 2.11.0+cu128
- **Datasets:** 2.21.0
- **Tokenizers:** 0.22.2