File size: 3,495 Bytes
8d9a29e
 
 
 
 
4705220
 
 
 
 
 
 
8d9a29e
 
 
 
 
 
 
 
 
 
4705220
8d9a29e
4705220
8d9a29e
4705220
 
 
 
8d9a29e
4705220
8d9a29e
4705220
 
 
 
 
8d9a29e
4705220
8d9a29e
4705220
8d9a29e
4705220
 
 
8d9a29e
4705220
 
 
8d9a29e
4705220
8d9a29e
4705220
 
 
8d9a29e
4705220
8d9a29e
4705220
 
 
 
 
 
 
 
 
8d9a29e
4705220
8d9a29e
4705220
8d9a29e
4705220
 
 
 
 
8d9a29e
4705220
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
---
library_name: transformers
license: apache-2.0
base_model: google/muril-base-cased
tags:
- ner
- token-classification
- indic
- multilingual
- marathi
- bengali
- telugu
metrics:
- precision
- recall
- f1
- accuracy
model-index:
- name: muril-ner-multilingual
  results: []
---

# muril-ner-multilingual 🏷️

This model is a fine-tuned version of **[google/muril-base-cased](https://huggingface.co/google/muril-base-cased)** for **Named Entity Recognition (NER)** across three major Indic languages: **Marathi (mr)**, **Bengali (bn)**, and **Telugu (te)**.

It uses a joint multilingual full fine-tuning strategy to extract three primary entity types:
- 👤 **PER** (Person)
- 🏢 **ORG** (Organization)
- 📍 **LOC** (Location)

## Model Description

- **Developed by:** Learner-sai
- **Model Type:** Token Classification (NER)
- **Base Model:** `google/muril-base-cased`
- **Languages:** Marathi (`mr`), Bengali (`bn`), Telugu (`te`)
- **License:** Apache 2.0

MuRIL (Multilingual Representations for Indian Languages) was adapted by adding a 7-class linear sequence classification head (`O`, `B-PER`, `I-PER`, `B-ORG`, `I-ORG`, `B-LOC`, `I-LOC`). All parameters were updated during joint training across all three target languages to leverage cross-lingual transfer.

## Intended Uses & Limitations

### Intended Uses 🎯
- Automated entity extraction from news articles, social media, and documents in Marathi, Bengali, and Telugu.
- Downstream NLP applications such as information retrieval, question answering, and knowledge graph construction for Indic languages.

### Limitations ⚠️
- **Grammatical Suffixes:** In agglutinative or highly inflected languages like Marathi and Telugu, locative or case suffixes (e.g., Marathi `-त` in "दिल्लीत") may sometimes be included inside the predicted entity span.
- **Entity Scope:** The model is trained exclusively on `PER`, `ORG`, and `LOC` tags; it will not recognize other categories like dates, monetary values, or product names.

## Training and Evaluation Data

The model was trained on a combined dataset comprising annotated sentences across Marathi, Bengali, and Telugu.
- **Label Schema:** IOB2 format (`B-`, `I-`, `O`) with 7 total classes.
- **Token Alignment:** Subword tokenization artifacts were handled using `-100` label masking on non-initial subwords to ensure clean cross-entropy loss calculation.

## Training Procedure

### Training Hyperparameters
- **Learning Rate:** `3e-05`
- **Train Batch Size:** 16
- **Eval Batch Size:** 32
- **Epochs:** 3
- **Optimizer:** `AdamW (fused)` with $\beta_1=0.9, \beta_2=0.999, \epsilon=1\text{e-}08$
- **LR Scheduler:** Linear with 937 warmup steps (~10% of total steps)
- **Mixed Precision:** Native FP16 (`fp16=True`)
- **Seed:** 42

### Evaluation Metrics

Evaluated on the validation split using `seqeval` (entity-level span matching):

| Epoch | Training Loss | Validation Loss | Precision | Recall | Entity F1 🏆 | Token Accuracy |
| :---: | :-----------: | :-------------: | :-------: | :----: | :---------: | :------------: |
| 1.0   | 0.2845        | 0.2822          | 0.7172    | 0.7725 | 0.7438      | 0.9320         |
| 2.0   | 0.2024        | 0.2335          | 0.7392    | 0.7723 | 0.7554      | 0.9346         |
| **3.0**| **0.1804**    | **0.2328**      | **0.7382**| **0.7771** | **0.7572**  | **0.9352**     |

## Framework Versions

- **Transformers:** 5.14.1
- **PyTorch:** 2.11.0+cu128
- **Datasets:** 2.21.0
- **Tokenizers:** 0.22.2