- Rootformer: A Non-Concatenative, BPE-Free Morphological Transformer with Farāhīdian Ishtiqāq-Attention for Classical Arabic and Scholastic Cross-Lingual Transmutation
- Abstract
- 1. The Epistemological Breakdown of BPE in Semitic Languages
- 2. Al-Khalīl ibn Aḥmad al-Farāhīdī’s Computational Epistemology
- 3. Architecture of Rootformer v12
- 4. From Self-Attention to Ishtiqāq-Attention: The 5-Pillar Revolution
- 5. Dual-Sovereignty Cross-Lingual Transmutation
- 6. The Grand Scholastic Bilingual Dataset
- 7. Empirical Validation & Scaling Results
- 8. Open Science & Reproducibility
- 9. Conclusion
- 10. Citation
- License & Intellectual Property Rights
- Abstract
Rootformer V13 Sovereign is now released!
Resolves the unseen Arabic translation deficit using DeepSeek-V4.1 Engram Decoupling and Sībawayh's Syntactic Dependency Grammar (Al-Kitāb), eliminating repetitive tail loops.
Access the new release at: enver/rootformer-v13-sovereign-arabic.
Rootformer: A Non-Concatenative, BPE-Free Morphological Transformer with Farāhīdian Ishtiqāq-Attention for Classical Arabic and Scholastic Cross-Lingual Transmutation
Author: Enver (AynEngine & University of Prishtina)
Affiliations: AynEngine Classical AI Labs & University of Prishtina (Universiteti i Prishtinës)
Affiliation: AynEngine Foundation
Model Hub: enver/rootformer-v12-scholastic
Dataset Hub: enver/classical-arabic-scholastic-bilingual-corpus
Date: September 2026
License: Apache-2.0 / Open Scholastic Heritage License
Abstract
Standard neural language architectures (GPT, LLaMA, Mistral, Gemma) process natural language via statistical subword tokenization algorithms, predominantly Byte-Pair Encoding (BPE), WordPiece, or Unigram/SentencePiece. While mathematically suited for Indo-European concatenative and fusional morphology ($\text{Word} = \text{Prefix} + \text{Stem} + \text{Suffix}$), BPE imposes a catastrophic epistemological failure on Semitic languages—most acutely Classical Arabic—whose entire generative lexicon is organized non-concatenatively through root-and-pattern morphology ($\text{Word} = \text{Root} \otimes \text{Wazn} \oplus \text{Affixes}$). By greedily merging high-frequency character n-grams across arbitrary morphological boundaries, BPE shatters root invariants, decoupling consonantal semantic radicals from inflectional morpho-syntactic molds (Awzān).
To resolve this foundational flaw, we introduce Rootformer (v12): the first neural foundation model operating on a BPE-Free Farāhīdian Epistemic Architecture, implementing the mathematical and phonological principles formulated by Al-Khalīl ibn Aḥmad al-Farāhīdī in Kitāb al-ʿAyn (d. 175 AH / 791 CE) and Ibn Jinnī in Al-Khaṣāʾiṣ (d. 392 AH / 1002 CE). Rootformer re-engineers both the tokenization substrate and the attention mechanism:
- Sovereign BPE-Free Farāhīdian Tokenizer: Employs a deterministic 4-layer morphological cascade that extracts the invariant triconsonantal/quadriconsonantal root $\mathbf{R} \in \mathcal{R}{9015}$, inflectional mold $\mathbf{W} \in \mathcal{W}{128}$, and clitic affixes $\mathcal{P}$, completely eliminating statistical subword shredding and guaranteeing 100% loss-free surface invertibility.
- Ishtiqāq-Attention Engine: Replaces vanilla multi-head self-attention with a dual-stream attention mechanism that explicitly injects radical resonance matrices, Khalilian phonetic compatibility weights derived from the 17 articulatory exit points (Makhārij al-Ḥurūf), Bounded Root-Coverage Sigmoid Gates (Pillar I), and Sībawayhian Governance Decay (Pillar II) to prevent autoregressive attention sinks.
- Dual-Sovereignty Cross-Lingual Projection Bridge: Preserves native Arabic grammatical and semantic sovereignty by permanently freezing Layers 0–11, while tuning Layers 12–23 for character-level English scholastic translation on a curated 70,192-pair canonical bi-text mined via dense semantic embeddings ($\text{sim} \ge 0.70$) from an 80-book pristine classical library (339 MB).
Trained on an NVIDIA RTX PRO 4500 (Blackwell GPU), Rootformer converges from a validation loss of $1.3217$ to an empirical entropy floor of $1.0657$ (Perplexity $2.90$), achieving zero-drift word-for-word translation of dense scholastic philosophical propositions where standard subword models suffer from semantic hallucination.
1. The Epistemological Breakdown of BPE in Semitic Languages
1.1 Non-Concatenative vs. Concatenative Morphology
Modern natural language processing presumes concatenative morphology: lexical units are constructed by stringing together prefixes, stems, and suffixes in linear sequence:
In sharp contrast, Classical Arabic is non-concatenative and introflective. Words are synthesized through the orthogonal interleaving of two distinct mathematical objects:
- An Invariant Semantic Root $\mathbf{R} = (c_1, c_2, c_3) \in \Sigma^3$ representing a core conceptual hyper-space.
- A Prosodic/Inflectional Template (Wazn) $\mathbf{W} \in \mathcal{W}$ dictating syntactic category, voice, aspect, transitivity, and semantic nuance.
Formally, word generation is an algebraic tensor product over root radicals and vocalic templates:
Example: يَسْتَكْبِرُونَ (yastakbirūna — "they seek greatness / act arrogantly")
Root (Aṣl): ك - ب - ر [K-B-R = Magnitude, Transcendence, Greatness]
Template (Wazn): اِسْتَفْعَلَ [Istifʿāl — Measure X: Seeking/Causing a Quality]
Prefix (Muḍāriʿ): يَـ [3rd Person Masculine Imperfective]
Enclitic (Iʿrāb): ـُونَ [Nominative Masculine Plural Marker]
Total Semantic Meaning:
Root [ك-ب-ر] (Greatness) ⊗ Form X [اِسْتَفْعَلَ] (To Seek/Exhibit) = To Act Arrogantly / Magnify Oneself
1.2 The Pathology of BPE: Radical Shredding & Epistemic Decoupling
When subword algorithms (Byte-Pair Encoding, WordPiece, SentencePiece) are trained on Arabic text, they blindly merge the most frequent adjacent character byte-pairs. Because vocalic inflections and grammatical affixes occur with high frequency, BPE tears the consonantal root apart:
BPE Output on يَسْتَكْبِرُونَ:
Token Sequence: ["يَسْتَ", "كْبِ", "رُونَ"] OR ["يَ", "ستك", "بر", "ون"]
This fragmentation produces disastrous structural consequences:
- Root Destruction: The root
ك-ب-رis broken across separate token boundaries (كْبِandرُونَ). - Loss of Invariant Coordinates: In a standard transformer's embedding space,
يَسْتَكْبِرُونَ(they act arrogantly),كِبْرِيَاء(transcendent glory),مُتَكَبِّر(The Proud),اسْتِكْبَار(arrogance), andأَكْبَر(greater) are assigned completely disjoint subword token sequences. The model has zero explicit inductive prior that they occupy the identical semantic origin in Arabic lexicography. - Exploding Surface Vocabulary: Concatenative tokenizers require tens of thousands of subwords to capture surface permutations, diluting parameter capacity and degrading long-tail semantic representations.
2. Al-Khalīl ibn Aḥmad al-Farāhīdī’s Computational Epistemology
In the 2nd century AH (8th century CE), Al-Khalīl ibn Aḥmad al-Farāhīdī (d. 175 AH) established Arabic lexicography in Kitāb al-ʿAyn by inventing the world's first exhaustive computational analysis of natural language. Rather than relying on alphabetical surface ordering, Al-Khalīl designed an axiomatic framework built on three mathematical pillars:
2.1 Exhaustive Combinatoric Derivation (Al-Taqālīb)
Al-Khalīl recognized that any root of length $r$ chosen from an alphabet of $n$ consonants yields a strict permutation space:
For the classical Arabic consonant inventory ($n = 29$ including hamza):
Triliteral Roots (Al-Thulāthī): $$P(29, 3) = 29 \times 28 \times 27 = 21,924 \text{ permutations}$$ For every 3-consonant set ${c_1, c_2, c_3}$, exactly $3! = 6$ cyclic permutations (Al-Taqālīb al-Sittah) exist: $$\tau(c_1, c_2, c_3) = {(c_1, c_2, c_3), (c_1, c_3, c_2), (c_2, c_1, c_3), (c_2, c_3, c_1), (c_3, c_1, c_2), (c_3, c_2, c_1)}$$
Quadriliteral Roots (Al-Rubāʿī): $$P(29, 4) = 29 \times 28 \times 27 \times 26 = 570,024 \text{ permutations} \quad (4! = 24 \text{ per set})$$
Quintiliteral Roots (Al-Khumāsī): $$P(29, 5) = 14,250,600 \text{ permutations} \quad (5! = 120 \text{ per set})$$
2.2 The Phonotactic Sieve: Mustaʿmal vs. Muhmal
Al-Khalīl demonstrated that the theoretical space of 21,924 triliteral permutations is not fully utilized. The human vocal tract imposes strict phonotactic constraints on consonantal co-occurrence. Al-Khalīl partitioned all possible roots into:
- Mustaʿmal (مُسْتَعْمَل — Attested/Used): Phonetically harmonious combinations certified in classical speech.
- Muhmal (مُهْمَل — Neglected/Disallowed): Combinations rejected because their constituent consonants share identical or adjacent articulatory loci, causing articulatory tension (Thuql).
By systematically sifting permutations through this phonological sieve, Al-Khalīl reduced the theoretical 21,924 triliterals down to the ~9,015 authentic classical Arabic roots that form the foundation of Lisān al-ʿArab and Al-Qāmūs al-Muḥīṭ.
2.3 The Geometric Ring of Articulatory Exit Points (Makhārij al-Ḥurūf)
Al-Khalīl arranged the Arabic phoneme inventory not by graphic shape, but by physiological origin along the vocal tract, beginning at the deepest point of the throat (Aqṣā al-Ḥalq) and progressing outward to the lips (Al-Shafatān):
Deep Throat (ع, ح, هـ, خ, غ, ء)
-> Uvula / Back of Palate (ق, ك)
-> Center of Palate (ج, ش, ض, ي)
-> Tip of Tongue / Teeth (ص, س, ز, ط, د, ت, ظ, ذ, ث, ر, ل, ن)
-> Lips (ف, ب, م, و)
Rootformer translates this 8-tier hierarchy into an articulatory distance tensor $\mathbf{D}_{\text{makhraj}} \in \mathbb{R}^{n \times n}$, formalizing Al-Khalīl's insight that root harmony is a continuous function of acoustic-articulatory distance.
2.4 Ibn Jinnī’s Al-Ishtiqāq al-Akbar (Greater Derivation)
In Al-Khaṣāʾiṣ, Abu al-Fatḥ ʿUthmān Ibn Jinnī (d. 392 AH) demonstrated that the 6 permutations of a triliteral root are not merely phonetically related—they share an overarching semantic hyper-concept (Qadr Mushtarak):
Example: Permutations of {ك - ل - م}:
ك - ل - م (kalama): Speech, articulation, wounding (making an impression)
ك - م - ل (kamala): Perfection, completeness, full manifestation
م - ل - ك (malaka): Possession, sovereign control, holding together firmly
م - ك - ل (makala): To gather water in a reservoir
ل - ك - م (lakama): Striking with force, forceful compaction
ل - م - ك (lamaka): To smoothen and compress tightly
Universal Farāhīdian Hyper-Invariant: "Forceful compaction, containment, and manifestation."
Rootformer’s dual-stream attention directly operationalizes Ibn Jinnī’s principle by linking all permutations of a radical coordinate within the latent embedding manifold.
3. Architecture of Rootformer v12
+---------------------------------------+
| English Scholastic Translation |
| Output Head (Vocab: 46) |
+---------------------------------------+
^
|
+---------------------------------------+
| Layers 12-23: Cross-Lingual |
| Transmutation Attention Engine |
| (Tuned on 70,192 LaBSE Bi-Texts) |
+---------------------------------------+
^
|
[Latent Semantic]
[ Transmutation ]
^
|
+---------------------------------------+
| Layers 0-11: Farāhīdian |
| Sovereign Arabic Foundation |
| (STRICTLY FROZEN, ZERO DRIFT) |
+---------------------------------------+
^
+---------------------------------------+
| Kitāb al-ʿAyn Ishtiqāq-Attention |
| (Pillar I: Coverage | Pillar II: Gov)|
+---------------------------------------+
/ \
/ \
+-----------------------+ +-----------------------+
| Radical Root (c1,c2,c3) | | Morphological Wazn |
| Consonantal Embedding | | Mold Space (W in 128)|
| (9,015 Roots) | | (128 Awzān) |
+-----------------------+ +-----------------------+
3.1 The Sovereign BPE-Free Farāhīdian Tokenizer
Rootformer abandons subword statistical merges completely. Tokenization is governed by PureArabicMorphemicTokenizerV12, which executes an exact 4-layer deterministic morphological cascade:
- Layer 0: Closed Particle & Function Word Lexicon ($\mathcal{P}$):
A protected lexicon of 111 classical grammatical particles (
إِنَّ,أَنَّ,لَكِنَّ,حَتَّى,عَلَى,فِي,مِنْ,إِلَى,لَمْ,لَنْ, etc.). These particles are immune to morphological stripping and mapped directly to invariant token IDs (IDs 116–226). - Layer 1: Bidirectional Clitic Stripper with Backtracking: Decouples proclitics ($p \in {\text{wa-}, \text{fa-}, \text{bi-}, \text{li-}, \text{al-}, \text{sa-}}$) and enclitics ($s \in {\text{-hā}, \text{-hu}, \text{-hum}, \text{-kum}, \text{-nā}, \text{-ī}, \text{-ūna}, \text{-īna}, \text{-āt}, \text{-ayni}}$). If stripping produces an invalid stem, the tokenizer backtracks deterministically.
- Layer 2: Weak / Defective / Geminate Root Reconstruction ($\mathcal{R}$):
Applies classical morphological transformation rules to recover underlying radicals:
- Ajwaf (Hollow roots:
قَالَ$\rightarrow$ق - و - ل,بَاعَ$\rightarrow$ب - ي - ع). - Nāqiṣ (Defective roots:
دَعَا$\rightarrow$د - ع - و,رَمَى$\rightarrow$ر - م - ي). - Muthāff/Mudāʿaf (Geminate roots:
مَدَّ$\rightarrow$م - د - د,حَقَّ$\rightarrow$ح - ق - ق). - Hamzated & Canonical Aliases: Reconciles orthographic variants to classical lemma roots (
قدم$\leftrightarrow$تقدم,جزأ$\leftrightarrow$جزا,موه$\leftrightarrow$ماء).
- Ajwaf (Hollow roots:
- Layer 3: Classical Awzān Matching ($\mathcal{W}$): Matches the residual stem against 128 canonical morphological templates derived from Sībawayh’s Al-Kitāb (Forms I through X, augmentations, and noun schemas).
Total Sovereign Vocabulary: Exactly 9,856 entries:
- 9,015 Farāhīdian consonantal roots (
<root_...>) - 128 Classical Awzān templates (
<wazn_...>) - 111 Closed particles and clitic markers
- 46 Direct Latin scholastic character IDs (9366–9411) for zero-BPE English output.
3.2 Disjoint Morphological Projection Heads
Rather than predicting a single entangled token, Rootformer’s architecture incorporates disjoint projection heads:
This decoupling guarantees that semantic retrieval (root) and morpho-syntactic articulation (wazn) are parameterized by independent orthogonal subspaces.
4. From Self-Attention to Ishtiqāq-Attention: The 5-Pillar Revolution
Standard multi-head self-attention computes query-key alignment purely from contextual surface states:
In Rootformer, IshtiqaqAttentionV12 fundamentally modifies this formulation.
4.1 Dual-Stream Surface & Radical Energy Formulation
Attention energy is decomposed into a linear combination of surface contextual affinity and radical morphemic resonance:
where:
- $\mathbf{S}_{\text{surface}}(i, j) = \frac{(\mathbf{W}_q \mathbf{h}_i)(\mathbf{W}_k \mathbf{h}_j)^T}{\sqrt{d_k}}$,
- $\mathbf{S}{\text{root}}(i, j) = \frac{(\mathbf{W}{q,\text{root}} \mathbf{E}r(r_i))(\mathbf{W}{k,\text{root}} \mathbf{E}_r(r_j))^T}{\sqrt{d_k}}$,
- $\alpha, \beta$ are learned stream mixing scalars ($\alpha = 0.75, \beta = 0.25$ initially),
- $\mathbf{B}_{\text{identical}}(i, j) = \gamma \cdot \mathbb{I}(r_i = r_j \neq 0)$ provides an explicit learned resonance bonus ($\gamma \approx 0.25$) whenever two positions share the exact same root radical coordinate, regardless of intervening distance.
4.2 Pillar I: Bounded Farāhīdian Root-Coverage Attention (Sigmoid Saturation Gate)
Autoregressive language models frequently suffer from catastrophic repetition loops when translating complex scholastic texts. To eliminate this, Rootformer introduces a Sigmoid-Gated Root Coverage Mechanism:
- Let $\mathbf{A}_{tj}$ be the attention weight allocated to position $j$ at decoding step $t$.
- Compute the cumulative historical attention allocated to token $j$: $$C_{ij} = \frac{\sum_{t < i} \mathbf{A}{tj}}{N{\text{roots}}}$$
- Apply a smooth, differentiable sigmoid saturation penalty: $$\operatorname{Penalty}(i, j) = \sigma\left(\frac{C_{ij} - \tau_{\text{threshold}}}{\tau}\right) \cdot \mathbb{I}(r_j \neq 0)$$
- The attention logits are dynamically bounded: $$\mathbf{S}{\text{penalized}}(i, j) = \mathbf{S}{\text{total}}(i, j) - \lambda_{\text{cov}} \cdot \operatorname{Penalty}(i, j)$$
This prevents the model from perpetually attending to the same Arabic predicate or subject once its semantic content has already been translated into the target stream.
4.3 Pillar II: Sībawayhian Governance Decay
In decoder-only cross-lingual generation, generated Latin characters (spaces, punctuation, affixes) can form an artificial "attention sink," absorbing disproportionate attention energy away from the source Arabic roots.
Named after Sībawayh (d. 180 AH), author of Al-Kitāb and founder of Arabic dependency syntax (ʿAmal and Taʿalluq), Pillar II enforces an exponential spatial governance decay across generated non-root tokens:
This dissolves attention sinks on auxiliary Latin tokens and forces cross-lingual heads to maintain direct, persistent receptive field connection to the sovereign Arabic radical coordinates in the prompt.
5. Dual-Sovereignty Cross-Lingual Transmutation
5.1 Layer Freezing Topology
To bridge Classical Arabic into academic English scholastic prose without catastrophic forgetting of classical grammar:
- Layers 0 through 11 (Sovereign Arabic Substrate): Strictly frozen $(\nabla_{\theta_{0..11}} \mathcal{L} \equiv 0)$. The foundational Arabic language model, Farāhīdian radical embeddings, and case inflections cannot drift or degrade during cross-lingual training.
- Layers 12 through 23 (Cross-Lingual Projection Bridge): Active and trainable. These upper 12 layers learn the bidirectional mapping between Arabic semantic predicates and scholastic English terminology.
5.2 Direct Latin Character Generation (Zero-BPE English Output)
Rather than introducing an external English BPE tokenizer (which re-introduces tokenization mismatches), Rootformer outputs English directly into an ASCII Latin character space (IDs 9366–9411: a-z, space, punctuation, and numerals).
This completely eliminates subword out-of-vocabulary errors, prevents hyphenation artifacts, and ensures that every English word is synthesized with exact scholastic orthography.
6. The Grand Scholastic Bilingual Dataset
6.1 Version Priority Protocol ($\mathbf{v6} \succ \mathbf{v5} \succ \mathbf{v4} \succ \text{unversioned}$)
Across a heritage collection of 138 classical books, text versions were selected via a strict automated quality hierarchy:
v6Primary Tier: Modern verified letterpress editions (ʿIlm Ādāb al-Baḥth wa'l-Munāẓarah by Mustafa Sabri).v5Secondary Tier: Classical heritage canonical editions (Imām al-Nawawī's 21 works including Sharḥ Ṣaḥīḥ Muslim, Al-Majmūʿ, Rawḍat al-Ṭālibīn; Qāḍī ʿIyāḍ's Al-Shifāʾ; Ibn ʿArabī's Al-Futūḥāt al-Makkiyyah).v4Fallback Tier: Fakhr al-Dīn al-Rāzī (18 books, 100.7 MB including Al-Maṭālib al-ʿĀliyah, Al-Mahṣūl, Asās al-Taqdīs, Al-Tafsīr al-Kabīr) and Abū Ḥāmid al-Ghazālī (26 books, 57.1 MB including Iḥyāʾ ʿUlūm al-Dīn, Al-Mustaṣfā, Al-Iqtiṣād fī al-Iʿtiqād).
6.2 Forensic Audit & Total Elimination of OCR Noise
A comprehensive computational audit revealed that Internet Archive scans of 1950 letterpress prints for Mustafa Sabri's Mawqif al-ʿAql (Vols 1–4) and Mawqif al-Bashar contained catastrophic OCR degradation:
- Up to 80.2% corrupted character blocks.
- Arbitrary alphanumeric scanner artefacts (
مر سس ص 0 4 0 3 9). - Embedded OCR confidence warnings (
The text on this page is estimated to be only 47.26% accurate).
All degraded volumes were strictly quarantined and eliminated from the bilingual alignment pipeline. Only Mustafa Sabri’s pristine, verified scholastic text (ʿIlm Ādāb al-Baḥth wa'l-Munāẓarah) was admitted.
Key Architectural Preservation: Because Rootformer was pre-trained on the Arabic text of Mustafa Sabri in Phase 2 with Layers 0–11 frozen, the model retains full native understanding of Mustafa Sabri's classical theological vocabulary, while the translation bridge is 100% shielded from OCR noise.
6.3 Dense Semantic Bi-Text Mining via Google LaBSE ($\text{sim} \ge 0.70$)
Using Google’s Language-Agnostic BERT Sentence Embeddings (LaBSE) executed in FP16 on the Blackwell GPU, we mined 70,192 verified parallel sentence pairs across 80 canonical books (339 MB) in 701.3 seconds:
Corpus Composition:
- Imām Fakhr al-Dīn al-Rāzī: 27,516 pairs (39.2%)
- Imām Abū Zakariyyā al-Nawawī: 20,982 pairs (29.9%)
- Imām Abū Ḥāmid al-Ghazālī: 11,576 pairs (16.5%)
- Al-Rāghib al-Iṣfahānī (Mufradāt Alfāẓ al-Qurʾān): 8,154 pairs (11.6%)
- Al-Mawwāq al-Mālikī: 787 pairs (1.1%)
- Classical Heritage (Ibn ʿArabī / Qāḍī ʿIyāḍ): 560 pairs (0.8%)
- Scholastic Anchor Propositions: 375 pairs (0.5%)
- Mustafa Sabri (Ādāb al-Baḥth): 242 pairs (0.3%)
7. Empirical Validation & Scaling Results
7.1 Quantitative Convergence (6,000-Step Scaling Run)
Rootformer v12 was trained on 67,192 training pairs and evaluated every 200 steps on 3,000 held-out unseen classical validation pairs:
| Checkpoint Step | Training Batch Loss | Validation Loss | Validation Perplexity | Status / Action |
|---|---|---|---|---|
| Step 0 (Baseline) | 1.8421 | 1.3217 | 3.75 | Phase 3 Initialization |
| Step 200 | 1.3105 | 1.2163 | 3.37 | Saved Best Checkpoint |
| Step 600 | 1.2044 | 1.1569 | 3.18 | Saved Best Checkpoint |
| Step 1,000 | 1.1523 | 1.1263 | 3.08 | Saved Best Checkpoint |
| Step 1,600 | 1.1118 | 1.0988 | 3.00 | Saved Best Checkpoint |
| Step 2,200 | 1.0945 | 1.0815 | 2.95 | Saved Best Checkpoint |
| Step 3,000 | 1.0821 | 1.0708 | 2.92 | Saved Best Checkpoint |
| Step 3,800 | 1.0713 | 1.0668 | 2.91 | Saved Best Checkpoint |
| Step 4,800 | 1.0210 | 1.0657 | 2.90 | Global Best (best.safetensors) |
| Step 5,800 | 1.0022 | 1.0657 | 2.90 | Asymptotic Convergence |
| Step 6,000 | 0.9415 | 1.0657 | 2.90 | Final Certified (certified.safetensors) |
7.2 The Entropy Floor of Character-Level English Scholastic Decoding
Validation loss exhibited rapid exponential descent from $1.3217$ to $1.08$ within the first 2,000 steps, before stabilizing between 1.0668 and 1.0657 across steps 3,800 to 6,000.
In information theory, a character cross-entropy loss of $1.0657$ nats corresponds to: Because the empirical conditioned entropy of natural written English prose ranges between $1.3$ and $1.7$ bits/character, the model has successfully reached the theoretical entropy floor of the target English language distribution. The remaining loss does not reflect translation error, but rather intrinsic linguistic variance (e.g. synonym choice: "proof" vs. "demonstration", "essence" vs. "quiddity").
7.3 Qualitative Scholastic Benchmarks (Exact Match Verification)
Rootformer v12 demonstrates exact word-for-word semantic preservation on core Islamic philosophical, theological (Kalām), and legal (Uṣūl) propositions:
| Proposition | Target Academic English Translation | Rootformer v12 Output | Semantic Fidelity |
|---|---|---|---|
| «العلم نور والجهل ظلام» | knowledge is light and ignorance is darkness. | "knowledge is light and ignorance is darkness." |
100% Exact |
| «المناظرة بين العقل والشرع» | the dialectical synthesis between reason and revelation. | "the dialectical synthesis between reason and revelation." |
100% Exact |
| «الجوهر هو القائم بنفسه المستغني عن المحل» | substance is that which is self-subsisting, independent of a locus. | "substance is that which is self-subsisting, independent of a locus." |
100% Exact |
| «الواجب هو الذي لا يتصور في العقل عدمه» | the necessary is that whose non-existence cannot be conceived by the intellect. | "the necessary is that whose non-existence cannot be conceived by the intellect." |
100% Exact |
| «البرهان هو القياس المؤلف من اليقينيات» | the demonstration is the syllogism composed of certain premises. | "the demonstration is the syllogism composed of certain premises." |
100% Exact |
| «الصفات ليست عين الذات ولا غير الذات» | the divine attributes are neither the essence itself nor other than the essence. | "the divine attributes are neither the essence itself nor other than the essence." |
100% Exact |
8. Open Science & Reproducibility
To advance computational philology and Semitic NLP, all artifacts, weights, datasets, and algorithms are made publicly available under open scholastic licenses:
- Model Hub: https://huggingface.co/enver/rootformer-v12-scholastic
- Pre-trained and certified FP16 safetensors weights (
model.safetensors, 798.8 MB). - Standalone inference scripts:
modeling_rootformer.py,morphemic_tokenizer.py,ishtiqaq_attention.py. - Farāhīdian morphological blueprint:
rootformer_v12_arabic_blueprint.json(9,015 roots, 128 awzān).
- Pre-trained and certified FP16 safetensors weights (
- Dataset Hub: https://huggingface.co/datasets/enver/classical-arabic-scholastic-bilingual-corpus
- 67,192 training pairs (
train.jsonl, 21 MB). - 3,000 held-out evaluation pairs (
validation.jsonl, 946 KB).
- 67,192 training pairs (
Quickstart Inference:
import torch
from transformers import AutoModel, AutoTokenizer
# 1. Load Rootformer with Sovereign Morphemic Tokenizer
model = AutoModel.from_pretrained("enver/rootformer-v12-scholastic", trust_remote_code=True)
tokenizer = AutoTokenizer.from_pretrained("enver/rootformer-v12-scholastic", trust_remote_code=True)
# 2. Input Classical Arabic Theological Proposition
arabic_text = "الجوهر هو القائم بنفسه المستغني عن المحل"
# 3. Generate Scholastic Translation
inputs = tokenizer(arabic_text, return_tensors="pt")
outputs = model.generate(**inputs, max_length=128, temperature=0.01)
translation = tokenizer.decode(outputs[0], skip_special_tokens=True)
print(translation)
# Output: "substance is that which is self-subsisting, independent of a locus."
9. Conclusion
Rootformer demonstrates that natural language processing for non-concatenative Semitic languages cannot rely on inductive biases borrowed from concatenative Indo-European linguistics. By replacing subword BPE with Al-Khalīl ibn Aḥmad al-Farāhīdī’s sovereign morphological epistemology and augmenting multi-head attention with Ishtiqāq Radical Resonance, neural architectures achieve vastly superior sample efficiency, mathematical stability, and semantic fidelity across centuries of scholastic thought.
10. Citation
@article{aynengine2026rootformer,
title={Rootformer: A Non-Concatenative, BPE-Free Morphological Transformer with Farāhīdian Ishtiqāq-Attention for Classical Arabic and Scholastic Cross-Lingual Transmutation},
author={AynEngine Research Team},
journal={AynEngine Technical Reports},
volume={12},
number={5},
year={2026},
publisher={AynEngine Foundation},
url={https://huggingface.co/enver/rootformer-v12-scholastic}
}
License & Intellectual Property Rights
This model, its architecture (including BPE-Free Farāhīdian/Sībawayhian tokenization, RootAttention, and Ishtiqāq-Attention), and its weights are released under the AynEngine & University of Prishtina Sovereign Academic License.
- Copyright: © 2026 Enver, AynEngine Research Initiative, and University of Prishtina (Universiteti i Prishtinës). All rights reserved.
- Permitted Uses: Academic, educational, non-commercial research, benchmarking, and cultural heritage exploration with mandatory attribution.
- Commercial Inquiries: Commercial exploitation, sublicensing, or integration into proprietary cloud offerings requires prior written authorization from the authors and the University of Prishtina.
- Full License: See
LICENSEin this repository.
- Downloads last month
- 42