File size: 6,412 Bytes
d0fadaa
 
b41a2d1
 
26a057b
fd5bfce
d0fadaa
b41a2d1
26a057b
b41a2d1
 
d0fadaa
b41a2d1
 
 
 
 
 
 
 
 
 
 
d0fadaa
 
1d7ea5a
7b6da35
1d7ea5a
7b6da35
1d7ea5a
10db1cf
7b6da35
 
26a057b
4cd565c
26a057b
 
 
b41a2d1
 
 
 
 
 
26a057b
50353fd
 
 
 
04e1769
b41a2d1
50353fd
 
 
b41a2d1
 
50353fd
 
 
 
 
b41a2d1
 
26a057b
 
 
 
 
d0fadaa
26a057b
 
b41a2d1
26a057b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b41a2d1
26a057b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b41a2d1
26a057b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
d0fadaa
26a057b
 
472dc39
26a057b
 
 
b41a2d1
 
 
 
 
26a057b
 
 
 
 
b41a2d1
 
 
d0fadaa
fd5bfce
d0fadaa
26a057b
 
b41a2d1
 
 
 
 
 
 
 
d0fadaa
fd5bfce
d0fadaa
b41a2d1
 
 
 
 
 
26a057b
 
 
 
 
b41a2d1
 
 
 
 
 
 
26a057b
b41a2d1
26a057b
 
 
b41a2d1
d0fadaa
26a057b
d0fadaa
b41a2d1
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
---
language:
  - ur

license: apache-2.0
library_name: transformers
pipeline_tag: fill-mask

datasets:
  - HPLT/HPLT3.0

tags:
  - modernbert
  - urdu
  - encoder
  - fill-mask
  - masked-language-modeling
  - long-context
  - 8k-context
  - urblimp
  - zero-shot
  - benchmark
  - safetensors
---

<h1 align="left">URNOVA</h1>

<div align="center">

<img src="./logo/logo.png" width="220">


</div>

URNOVA is a compact Urdu masked-language model encoder built with the ModernBERT architecture.

It was trained from scratch on an Urdu corpus derived from HPLT 3.0 and is designed for:

- Urdu representation learning
- masked-language modelling
- semantic feature extraction
- document and sentence classification
- downstream Urdu NLP fine-tuning
- research involving longer Urdu sequences

## Benchmark results

### UrBLiMP zero-shot linguistic evaluation

Zero-shot minimal-pair evaluation on UrBLiMP using left-to-right word pseudo-log-likelihood (`pll_word_l2r`). 


| Model | Params | Pairs | Micro Acc. | Macro | Aspect Agreement | Dative Object | Ergativity | Experiencer Subject | Honorific | Noun Phrase | Oblique | Participial Relatives | Subject-Verb Agreement | Word Order |
| --- | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: | ---: |
| **URNOVA-95M** | **95.2M** | 5,698 | 98.17% | 97.96% | **100.00%** | 96.23% | 99.10% | **98.77%** | **97.39%** | **99.50%** | 98.22% | 94.68% | 96.66% | **99.01%** |
| HPLT-BERT Urdu | 150.2M | 5,698 | **98.48%** | **98.14%** | 99.13% | **99.14%** | **99.28%** | 97.04% | 96.41% | 97.75% | **98.54%** | **98.01%** | **97.62%** | 98.51% |
| XLM-R-large | 560.1M | 5,698 | 95.75% | 94.31% | 99.25% | 96.23% | 96.20% | 95.31% | 79.74% | 93.00% | 96.59% | 92.36% | 96.40% | 98.02% |
| Urdu-RoBERTa-small | 126.0M | 5,698 | 94.02% | 90.75% | 99.75% | 88.70% | 97.05% | 95.80% | 70.59% | 92.50% | 96.23% | 89.70% | 89.07% | 88.12% |
| XLM-R-base | 278.3M | 5,698 | 92.70% | 89.32% | 99.50% | 93.84% | 95.36% | 92.84% | 61.44% | 88.50% | 96.52% | 79.40% | 87.79% | 98.02% |
| mmBERT-small | 140.9M | 5,698 | 91.08% | 88.33% | 99.25% | 91.78% | 95.06% | 73.83% | 65.36% | 87.00% | 92.89% | 89.70% | 89.46% | **99.01%** |


## Quick start

### Fill-mask inference

```python
from transformers import pipeline

fill_mask = pipeline(
    task="fill-mask",
    model="ProximaAI/urnova-95m",
    token=True,
)

results = fill_mask(
    "پاکستان کی قومی زبان [MASK] ہے۔",
    top_k=5,
)

for result in results:
    print(
        result["token_str"],
        result["score"],
    )
```

### Load the masked-language model

```python
from transformers import (
    AutoModelForMaskedLM,
    AutoTokenizer,
)

model_id = "ProximaAI/urnova-95m"

tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    token=True,
)

model = AutoModelForMaskedLM.from_pretrained(
    model_id,
    token=True,
)
```

### Extract contextual representations

```python
import torch
from transformers import AutoModel, AutoTokenizer

model_id = "ProximaAI/urnova-95m"

tokenizer = AutoTokenizer.from_pretrained(
    model_id,
    token=True,
)

encoder = AutoModel.from_pretrained(
    model_id,
    token=True,
)

text = "اردو ایک خوب صورت اور وسیع زبان ہے۔"

inputs = tokenizer(
    text,
    return_tensors="pt",
    truncation=True,
    max_length=8192,
)

with torch.inference_mode():
    output = encoder(**inputs)

token_embeddings = output.last_hidden_state
attention_mask = inputs["attention_mask"].unsqueeze(-1)

sentence_embedding = (
    token_embeddings * attention_mask
).sum(dim=1) / attention_mask.sum(dim=1)
```

## Training data and provenance

URNOVA was trained from scratch using an Urdu corpus derived from HPLT 3.0.

Recorded corpus details:

- 1,921,027 source documents
- 1,246,973,570 source tokens
- approximately 4 billion total training tokens
- tokenizer vocabulary size of 50,048
- no third-party model initialization

No raw training corpus is distributed in this repository.

See:

- [`DATA_AND_THIRD_PARTY_NOTICE.md`](./DATA_AND_THIRD_PARTY_NOTICE.md)
- [`NOTICE`](./NOTICE)
- [`LICENSE-WEIGHTS`](./LICENSE-WEIGHTS)

## Intended use

URNOVA is intended for:

- Urdu NLP research
- masked-language modelling
- encoder-based feature extraction
- document classification
- sentence classification
- semantic similarity research
- downstream supervised fine-tuning
- experimentation with long Urdu documents

## Limitations

- The reported benchmark evaluates grammatical minimal-pair preference on UrBLiMP and should not be interpreted as a complete measure of general Urdu language understanding.
- Benchmark performance may not generalize to Roman Urdu, code-switched text, regional dialects, or unrelated domains.
- Minimal-pair pseudo-log-likelihood evaluation measures sentence preference rather than generative ability, reasoning ability, or downstream task performance.
- UrBLiMP results depend on the evaluation and scoring procedure described above and should only be compared with results produced using a compatible methodology.
- Long-context capability does not guarantee reliable understanding over every 8,192-token input.
- Training on web-derived text may reproduce factual errors, stereotypes, offensive content, or representation gaps found in the source material.

## Responsible use

Predictions from URNOVA should not be the sole basis for:

- punitive decisions
- political profiling
- content removal
- employment decisions
- legal decisions
- eligibility decisions
- other high-impact automated actions

Applications should include domain-specific validation, human review, and appropriate error analysis.

## License and data notice

The original URNOVA model weights, tokenizer artifacts, code, documentation, benchmark reports, and visualizations are made available under the **Apache License 2.0**, to the extent of the rights held by their respective copyright holders.

The model was trained using an Urdu corpus derived from HPLT 3.0. HPLT licenses its dataset packaging under CC0 but states that it does not own the underlying extracted web text.

Apache-2.0 therefore does not grant rights in underlying training documents or third-party text that may be reproduced by the model.

Third-party benchmark datasets are not covered by URNOVA's Apache-2.0 license. UrBLiMP remains subject to its upstream MIT License and copyright notice.