ruqlm / README.md
Ruqiya's picture
Self-contained PyTorch example, escaped regex ranges
42e796c verified
|
Raw
History Blame Contribute Delete
8.96 kB
---
license: apache-2.0
language: [ar]
library_name: onnx
pipeline_tag: text-generation
tags: [arabic, story-generation, scaling-laws, onnx, small-language-model]
---
# RuqLM โ€” small Arabic language models
**โ–ถ Try it live: [ruqlm.web.app](https://ruqlm.web.app)** โ€” runs entirely in your
browser, no server, no signup needed for guest access.
A ladder of five Arabic language models, **0.74M to 29.89M parameters**, trained
from scratch on a corpus of **50,003 Arabic stories** under a fixed budget of 76.8M
training tokens. The tokenizer, the corpus and the weights were all built for this
project.
Weights are **full-precision ONNX** โ€” no quantisation. Perplexity matches the
original PyTorch checkpoints to three decimals (9.2009 vs 9.2009), so these are the
trained models, not an approximation of them.
---
## The ladder
| Rung | Parameters | d_model | Layers | Heads | `.onnx` | `.safetensors` |
|---|---|---|---|---|---|---|
| `ruq-30m` | 29.89M | 512 | 8 | 8 | 114.8 MB | 114.0 MB |
| `ruq-15m` | 13.77M | 384 | 6 | 6 | 53.1 MB | 52.5 MB |
| `ruq-5m` | 5.31M | 256 | 4 | 4 | 20.7 MB | 20.3 MB |
| `ruq-2m` | 1.90M | 128 | 4 | 4 | 7.7 MB | 7.3 MB |
| `ruq-0.7m` | 0.74M | 64 | 4 | 4 | 3.2 MB | 2.8 MB |
Decoder-only transformer: RMSNorm, RoPE, SwiGLU, tied embeddings, 512-token context.
Tokenizer is a byte-level BPE with an 8,192 vocabulary trained on the same corpus
(1.411 tokens per word).
---
## Usage
```bash
pip install onnxruntime tokenizers huggingface_hub numpy
```
Colab has none of these preinstalled except `numpy`, so run the line above first
(prefix it with `!` in a notebook cell).
```python
import re, unicodedata
import numpy as np, onnxruntime as ort
from huggingface_hub import hf_hub_download
from tokenizers import Tokenizer
REPO, EOS = "Ruqiya/ruqlm", 2
tok = Tokenizer.from_file(hf_hub_download(REPO, "tokenizer.json"))
sess = ort.InferenceSession(hf_hub_download(REPO, "ruq-30m.onnx"))
# The tokenizer was trained on normalised text: diacritics and tatweel
# removed. Skipping this step gives a different, worse tokenisation.
_DIACRITICS = re.compile(r"[\u064b-\u0652\u0670\u0653-\u0655]")
_ZERO_WIDTH = re.compile(r"[\u200b-\u200f\u202a-\u202e\ufeff]")
def normalize(text):
text = unicodedata.normalize("NFC", text)
text = _ZERO_WIDTH.sub("", text).replace("ู€", "")
return re.sub(r"\s+", " ", _DIACRITICS.sub("", text)).strip()
def generate(prompt, max_new_tokens=120, temperature=0.85, top_k=50, seed=None):
"""The model outputs logits for the last position only; sampling is here."""
rng = np.random.default_rng(seed)
ids = tok.encode(normalize(prompt)).ids
if ids and ids[-1] == EOS: # drop the trailing </s> so it continues
ids = ids[:-1]
for _ in range(max_new_tokens):
logits = sess.run(None, {"input_ids": np.array([ids], dtype=np.int64)})[0][0]
logits = logits.astype(np.float64) / temperature
kth = np.partition(logits, -top_k)[-top_k] # top-k filter
logits[logits < kth] = -np.inf
probs = np.exp(logits - logits.max())
probs /= probs.sum()
nxt = int(rng.choice(len(probs), p=probs))
if nxt == EOS:
break
ids.append(nxt)
return tok.decode(ids)
print(generate("ูƒุงู† ูŠุง ู…ุง ูƒุงู†"))
```
Output (`ruq-5m`, seed 0):
> ูƒุงู† ูŠุง ู…ุง ูƒุงู† ููŠ ุญุฏูŠู‚ุฉ ุจูŠุช. ูƒุงู† ุงู„ู†ุฌุงุฑ ูŠุนู…ู„ ุจุฌุฏ ู„ูŠุทุนู… ุนุงุฆู„ุชู‡. ูˆููŠ ูŠูˆู… ู…ู† ุงู„ุฃูŠุงู…ุŒ
> ูˆุฃุซู†ุงุก ุนู…ู„ู‡ุŒ ูˆุฌุฏ ุงู„ู†ุฌุงุฑ ุณู„ุญูุงุฉ ุตุบูŠุฑุฉ ุถุงุฆุนุฉโ€ฆ
### PyTorch
The same weights are published as safetensors alongside the architecture module, for
continued training or fine-tuning:
```python
import json, re, unicodedata, torch
from huggingface_hub import hf_hub_download
from safetensors.torch import load_file
from tokenizers import Tokenizer
REPO, EOS = "Ruqiya/ruqlm", 2
tok = Tokenizer.from_file(hf_hub_download(REPO, "tokenizer.json"))
_DIACRITICS = re.compile(r"[\u064b-\u0652\u0670\u0653-\u0655]")
_ZERO_WIDTH = re.compile(r"[\u200b-\u200f\u202a-\u202e\ufeff]")
def normalize(text):
text = unicodedata.normalize("NFC", text)
text = _ZERO_WIDTH.sub("", text).replace("\u0640", "") # tatweel
return re.sub(r"\s+", " ", _DIACRITICS.sub("", text)).strip()
# the architecture module has to be on the path before it can be imported
hf_hub_download(REPO, "modeling_ruqlm.py", local_dir=".")
from modeling_ruqlm import RuqLM, ModelArgs
cfg = json.load(open(hf_hub_download(REPO, "configs.json")))["ruq-5m"]
model = RuqLM(ModelArgs(**cfg))
state = load_file(hf_hub_download(REPO, "ruq-5m.safetensors"))
state["lm_head.weight"] = state["tok_emb.weight"] # embeddings are tied
model.load_state_dict(state)
model.eval()
ids = tok.encode(normalize("ูƒุงู† ูŠุง ู…ุง ูƒุงู†")).ids[:-1] # drop the trailing </s>
out = model.generate(torch.tensor([ids]), max_new_tokens=120,
temperature=0.85, top_k=50, eos_id=EOS)
print(tok.decode(out[0].tolist()))
```
`lm_head` and `tok_emb` are the same tensor, so only one is stored and the tie is
restored on load. `modeling_ruqlm.py` depends on nothing but `torch`.
### Interface
The graph takes `input_ids` of shape `[1, seq]` and returns `logits` for the **last
position only**, shape `[1, 8192]`. The sampling loop lives outside the model, so
temperature and top-k can change without re-exporting.
### In the browser
```html
<script src="https://cdn.jsdelivr.net/npm/onnxruntime-web/dist/ort.min.js"></script>
<script type="module">
import { AutoTokenizer } from
'https://cdn.jsdelivr.net/npm/@huggingface/transformers/+esm';
const REPO = 'https://huggingface.co/Ruqiya/ruqlm/resolve/main';
const tok = await AutoTokenizer.from_pretrained('Ruqiya/ruqlm');
const sess = await ort.InferenceSession.create(`${REPO}/ruq-5m.onnx`);
let ids = Array.from((await tok('ูƒุงู† ูŠุง ู…ุง ูƒุงู†')).input_ids.data, Number);
if (ids.at(-1) === 2) ids = ids.slice(0, -1); // drop </s>
for (let i = 0; i < 100; i++) {
const input = new ort.Tensor('int64', BigInt64Array.from(ids, BigInt),
[1, ids.length]);
const { logits } = await sess.run({ input_ids: input });
// greedy for brevity; sample with temperature and top-k for better text
const next = logits.data.indexOf(Math.max(...logits.data));
if (next === 2) break;
ids.push(next);
}
console.log(tok.decode(ids, { skip_special_tokens: true }));
</script>
```
The tokenizer normalises text the same way the Python example does, so no separate
normalisation step is needed here.
### Fine-tuning
The safetensors carry weights only โ€” no optimiser state โ€” so a fine-tune starts
from a fresh optimiser:
```python
model.train()
opt = torch.optim.AdamW(model.parameters(), lr=1e-4)
# input_ids: LongTensor [batch, seq]; the model shifts internally, so labels
# are the inputs themselves
logits, loss = model(input_ids, labels=input_ids)
loss.backward()
torch.nn.utils.clip_grad_norm_(model.parameters(), 1.0)
opt.step(); opt.zero_grad()
```
Context length is 512. Chat-role tokens `<|user|>` (id 4) and `<|assistant|>`
(id 5) are already reserved in the vocabulary, so instruction tuning needs no
tokenizer change.
---
## Limitations
These are **base text-completion models, not chat models.** They were trained on
Arabic stories only, so they continue what you start; they do not answer questions or
follow instructions, having had no instruction tuning (SFT) or RLHF. Expect
narrative, not answers.
Morphology and orthography are strong. **Semantic and referential coherence are
weaker** โ€” which is what the accompanying study measures, and what these models were
built to quantify rather than to solve. Even the largest rung abandons 5.9% of the
characters it introduces, where the training corpus abandons 0.0%.
The corpus is synthetic, generated by two teacher models: **ALLaM 2.7B** (75%) and
**Qwen3.6-27B** (25%), with quality gates rejecting foreign characters and template
openings. Both teachers are Apache-2.0.
---
## What the models are for
They were built to answer a measurement question: across a 40ร— range of parameters
at fixed data, which linguistic competence scales and which saturates?
| Axis | Behaviour across the ladder |
|---|---|
| Lexical diversity (TTR) | **saturated** โ€” 0.719 ยฑ 0.005 โ†’ 0.721 ยฑ 0.005 (0.4ฯƒ) |
| Character tracking | **scales** โ€” 31.6% โ†’ 5.9% abandoned (6.3ฯƒ) |
| Orthography, agreement | saturated below 0.74M |
Capacity buys coherence, not vocabulary.
---
## Paper
<!-- TODO: arXiv link on release -->
*In preparation.* The paper covers the corpus construction, the evaluation
framework, and every measurement reported here.
```bibtex
@misc{binsafi2026ruqlm,
title = {What Scales in a Small Arabic Language Model?
Morphology Saturates, Coherence Does Not},
author = {Bin Safi, Ruqiya},
year = {2026},
note = {Preprint}
}
```