File size: 3,473 Bytes
f91db99
ee64f10
 
 
f91db99
ee64f10
 
 
 
 
 
 
 
 
f91db99
 
ee64f10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
684397b
ee64f10
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
---
language:
- ar
license: apache-2.0
tags:
- mamba
- state-space-model
- arabic
- masked-language-modeling
- bidirectional
datasets:
- wikimedia/wikipedia
- uonlp/CulturaX
pipeline_tag: fill-mask
---

# AraSSM-base

AraSSM is a bidirectional state-space (Mamba) encoder pretrained from scratch for Arabic via
masked language modeling. It is, to our knowledge, the first bidirectional Mamba/SSM encoder
pretrained specifically for Arabic, and was trained entirely on four consumer-grade NVIDIA RTX
2080Ti GPUs (11GB) rather than an accelerator cluster.

## Model description

Each AraSSM layer runs a forward and a backward selective-scan (Mamba) mixer over the same
input and merges the two outputs, giving the model full bidirectional context while keeping
$O(L)$ complexity in sequence length $L$, instead of the $O(L^2)$ complexity of self-attention.

| | |
|---|---|
| Layers | 12 |
| Hidden size  | 512 |
| State dimension | 16 |
| Parameters | ~105M |
| Max sequence length | 512 |
| Vocabulary size | 64,000 |

## Tokenizer

AraSSM uses the existing **AraBERTv02 tokenizer**
([`aubmindlab/bert-base-arabertv02`](https://huggingface.co/aubmindlab/bert-base-arabertv02))
rather than a custom-trained one. This tokenizer was not trained as part of this work; it is
reused as-is, both to decouple corpus preprocessing from tokenizer choice and to keep results
comparable to AraBERT-family baselines that use the same vocabulary. All credit for the
tokenizer belongs to its original authors.

## Training data

AraSSM is pretrained on a cleaned, deduplicated corpus of approximately 80GB of Arabic text
(79.6GB train / 0.4GB validation), combining Arabic Wikipedia and the Arabic portion of
CulturaX. Documents are stripped of diacritics and URLs, filtered for length and Arabic-script
ratio, chunked to at most 400 words, and deduplicated at the chunk level with a Bloom filter.

## Training procedure

- Objective: standard BERT-style masked language modeling (15% masking, 80/10/10 split)
- Optimizer: AdamW, lr 3e-4, weight decay 0.01, 10,000 warmup steps, linear decay
- Precision: fp16 (Turing GPUs do not support accelerated bf16)
- Effective batch size: 256 (batch size 8 x grad accumulation 8 x 4 GPUs)
- Compute: 4x RTX 2080Ti, ~960 GPU-hours (~10 days)

## How to use

AraSSM is a custom architecture (not a native `transformers` model class), so loading it
requires the model code from the project repository:

```python
from huggingface_hub import PyTorchModelHubMixin
from models.mamba import MambaForMaskedLM  # from the AraSSM repository

class HubMambaForMaskedLM(MambaForMaskedLM, PyTorchModelHubMixin):
    pass

model = HubMambaForMaskedLM.from_pretrained("aliane29/arassm-base")

from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("aliane29/arassm-base")
```

## Intended use

This is a pretrained encoder intended to be fine-tuned on downstream Arabic NLU tasks
(classification, token classification, extractive question answering), similarly to how a
BERT-family encoder is used. It has not been fine-tuned for any specific task in this repository.

## Limitations

- Pretrained on Modern Standard Arabic and web text (Wikipedia + CulturaX); performance on
  dialectal Arabic is not evaluated.
- Maximum sequence length is 512 tokens.
- Trained on a fixed, publicly available compute budget (4 consumer GPUs); larger-scale
  Transformer baselines were pretrained on substantially larger accelerator-cluster budgets.