Mirrored from github.com/SNAPKITTYWEST/symbolic-morphology at commit
d23a58c.
✅ Independently reproduced (2026-10-07): every accuracy and loss figure for Rust and Python matched the README exactly. Rust 96/96 train and 83.2% unseen, final loss 9.1009e-5 on both engines, NAND 83.8% unseen. Details: VERIFICATION.md.
Symbolic Morphology Engine
A from-scratch symbolic learning engine that acquires Latin verb morphology from raw letter sequences and converges toward Boolean grammatical representations — implemented in six runtimes, from NAND gates to F#.
Table of Contents
- Overview
- Architecture
- Mathematical Core
- Repository Structure
- Building and Running
- Benchmarks
- Attention Kernel
- The Three Layers
- Training Corpus
- Boolean Convergence
- Generalization
- The NAND Variant
- Extending the System
- Mathematical Reference
- License
Overview
This project implements a complete neural morphology learner from mathematical primitives upward. No PyTorch, no TensorFlow, no scikit-learn, no automatic differentiation. Every derivative is computed by hand using the classical chain rule. Every arithmetic operation is implemented from scratch.
The system learns to map Latin verb forms — raw sequences of characters like AMO, AMABAT, REGUNT — to Boolean grammatical feature vectors encoding person, number, tense, mood, voice, and conjugation class. It does this by:
- Representing each character as a learnable embedding vector
- Processing the character sequence through a differentiable computation graph
- Producing continuous outputs in [0, 1] via sigmoid activation
- Thresholding at 0.5 to obtain Boolean grammatical predictions
- Computing binary cross-entropy loss against known targets
- Propagating gradients backward through the chain rule
- Updating parameters via gradient descent
The same machinery operates on every training example. No word is hardcoded. No rule is hand-specified. The system discovers statistical regularities in character sequences — stems, suffixes, positional patterns — and maps them to grammatical categories through gradient-based optimization.
Six implementations are provided, spanning the full spectrum from silicon-level abstraction to functional programming:
| Runtime | Language | Paradigm | Dependencies |
|---|---|---|---|
src/rust |
Rust | Systems, zero-cost abstractions | Zero |
src/csharp |
C# | OOP, imperative | .NET 8 BCL only |
src/fsharp |
F# | Functional, expression-oriented | .NET 8 BCL only |
src/python (float) |
Python | Scripting, float arithmetic | Python stdlib only |
src/python (NAND) |
Python | Gate-level simulation | Python stdlib only |
Each implementation is self-contained and independently runnable.
Architecture
The system follows a strict three-layer separation inspired by the observation/learning/symbolic distinction in cognitive architectures.
flowchart TB
subgraph OBS["OBSERVATION LAYER"]
direction LR
L1["Latin Word<br/>AMO"] --> L2["Character Array<br/>['A','M','O']"]
L2 --> L3["Character Indices<br/>[65, 77, 79]"]
end
subgraph LEARN["LEARNING LAYER"]
direction LR
E["Embedding Lookup<br/>e_A, e_M, e_O"] --> P["Mean Pooling<br/>h = Σe_t / N"]
P --> H["Hidden Layer<br/>a₁ = tanh(W₁h + b₁)"]
H --> O["Output Layer<br/>ŷ = σ(W₂a₁ + b₂)"]
end
subgraph SYM["SYMBOLIC LAYER"]
direction LR
O2["Continuous Output<br/>[0.97, 0.02, 0.01, ...]"] --> T["Threshold at 0.5"]
T --> B["Boolean Features<br/>PERSON_1=TRUE<br/>SINGULAR=TRUE<br/>PRESENT=TRUE"]
end
OBS --> LEARN --> SYM
style OBS fill:#1a1a2e,stroke:#e94560,color:#fff
style LEARN fill:#16213e,stroke:#0f3460,color:#fff
style SYM fill:#1a1a2e,stroke:#533483,color:#fff
Data Flow
flowchart LR
W["Word:<br/>AMO"] --> C["Characters:<br/>A, M, O"]
C --> E["Embeddings:<br/>e₆₅, e₇₇, e₇₉<br/>(each 16-dim)"]
E --> MP["Mean Pool:<br/>h = (e₆₅+e₇₇+e₇₉)/3<br/>(16-dim vector)"]
MP --> HL["Hidden Layer:<br/>z₁ = W₁h + b₁<br/>a₁ = tanh(z₁)<br/>(32-dim vector)"]
HL --> OL["Output Layer:<br/>z₂ = W₂a₁ + b₂<br/>ŷ = σ(z₂)<br/>(23-dim vector)"]
OL --> TH["Threshold:<br/>ŷᵢ ≥ 0.5 → TRUE<br/>ŷᵢ < 0.5 → FALSE"]
TH --> F["Grammatical<br/>Features"]
style W fill:#2d2d44,color:#fff
style F fill:#2d2d44,color:#fff
Mathematical Core
Every implementation computes the same mathematical operations. Nothing is hidden behind a framework.
Forward Pass
For input word w = c₁c₂...cₙ (sequence of characters):
1. Embedding lookup:
e_t = Embeddings[c_t] for t = 1..N
2. Mean pooling (variable-length → fixed-size):
h = (1/N) Σ e_t
3. Hidden layer (tanh activation):
z₁ = W₁h + b₁
a₁ = tanh(z₁)
4. Output layer (sigmoid activation):
z₂ = W₂a₁ + b₂
ŷ = σ(z₂)
Loss Function
Binary cross-entropy over all K output features:
L = -(1/K) Σᵢ [ yᵢ log(ŷᵢ) + (1-yᵢ) log(1-ŷᵢ) ]
Backward Pass — Classical Chain Rule
Every derivative is computed using the local derivative rule:
gradient_in = gradient_out × local_derivative
At branching points (where one value feeds multiple downstream operations):
gradient = Σ(all incoming gradient contributions)
The full backward chain:
flowchart BT
L["Loss L"] --> |"∂L/∂z₂ = (ŷ-y)/K"| DZ2["dL/dz₂"]
DZ2 --> |"∂z₂/∂W₂ = a₁"| GW2["dL/dW₂ = dL/dz₂ · a₁ᵀ"]
DZ2 --> |"∂z₂/∂b₂ = 1"| GB2["dL/db₂ = dL/dz₂"]
DZ2 --> |"∂z₂/∂a₁ = W₂"| DA1["dL/da₁ = W₂ᵀ · dL/dz₂"]
DA1 --> |"∂a₁/∂z₁ = 1-a₁²"| DZ1["dL/dz₁ = dL/da₁ ⊙ (1-a₁²)"]
DZ1 --> |"∂z₁/∂W₁ = h"| GW1["dL/dW₁ = dL/dz₁ · hᵀ"]
DZ1 --> |"∂z₁/∂b₁ = 1"| GB1["dL/db₁ = dL/dz₁"]
DZ1 --> |"∂z₁/∂h = W₁"| DH["dL/dh = W₁ᵀ · dL/dz₁"]
DH --> |"∂h/∂e_t = 1/N"| DE["dL/de_t = (1/N) · dL/dh"]
style L fill:#e94560,color:#fff
style GW2 fill:#0f3460,color:#fff
style GB2 fill:#0f3460,color:#fff
style GW1 fill:#0f3460,color:#fff
style GB1 fill:#0f3460,color:#fff
style DE fill:#533483,color:#fff
Parameter Update
θ := θ - η ∇θL
where η = learning rate (0.5 in our experiments)
Hyperparameters
| Parameter | Value | Rationale |
|---|---|---|
| Embedding dimension | 16 | Enough to distinguish ~20 Latin characters |
| Hidden dimension | 32 | Provides nonlinear feature combinations |
| Learning rate | 0.5 | Aggressive; loss landscape is smooth at this scale |
| Epochs | 3,000 | Converges well before this; ensures full training |
| Weight initialization | Xavier/Glorot | scale = sqrt(2 / (fan_in + fan_out)) |
| PRNG seed | 42 | Reproducible across all compiled implementations |
Repository Structure
symbolic-morphology/
├── README.md # This file (~5,000 words)
├── LICENSE # GNU Affero General Public License v3
├── .gitignore # Build artifacts, IDE files
│
├── src/
│ ├── rust/ # Rust implementation (zero dependencies)
│ │ ├── Cargo.toml # Package manifest
│ │ ├── src/
│ │ │ ├── main.rs # Entry point, benchmark harness
│ │ │ ├── engine.rs # Forward/backward/gradient descent
│ │ │ ├── features.rs # Boolean feature schema
│ │ │ ├── dataset.rs # Latin verb corpus
│ │ │ ├── lib.rs # Library root
│ │ │ ├── attention.rs # Attention kernel
│ │ │ └── bin/
│ │ │ ├── attention.rs # Attention benchmark
│ │ │ └── lean.rs # Training loop (pool=attn | mean)
│ │ └── tests/
│ │ └── lean_cli.rs # lean binary integration tests
│ │
│ ├── csharp/ # C# implementation (.NET 8)
│ │ ├── Sovereign.Engine.csproj # Project file
│ │ ├── Program.cs # Entry point, benchmark harness
│ │ ├── Engine.cs # Forward/backward/gradient descent
│ │ ├── Features.cs # Boolean feature schema
│ │ └── Dataset.cs # Latin verb corpus
│ │
│ ├── fsharp/ # F# implementation (.NET 8)
│ │ ├── Sovereign.Engine.FSharp.fsproj
│ │ ├── Attention.fs # Attention kernel
│ │ ├── AttentionTests.fs # Attention tests and benchmark
│ │ └── Program.fs # Complete engine + benchmarks
│ │
│ └── python/ # Python implementations
│ ├── symbolic_bench.py # Float-based Elman RNN
│ ├── nand_latin.py # NAND-gate recursive variant
│ └── attention.py # Attention kernel
│
├── mojo/
│ ├── morphology.mojo # Mojo engine (pool=mean | attn)
│ ├── attention.mojo # Attention kernel
│ ├── attention_test.mojo # Attention tests
│ └── attention_bench.mojo # Attention benchmark
│
├── docs/
│ ├── ARCHITECTURE.md # Deep-dive architecture document
│ ├── BENCHMARKS.md # Full benchmark results and analysis
│ └── ATTENTION.md # Attention kernel documentation
│
└── bench/
├── results.md # Comparative benchmark table
└── shared/
└── attention_golden.txt # Attention golden vectors
Building and Running
Prerequisites
| Implementation | Requirement |
|---|---|
| Rust | rustc 1.98+ / cargo 1.98+ |
| C# | .NET 8 SDK |
| F# | .NET 8 SDK (includes F# compiler) |
| Python | Python 3.12+ (stdlib only) |
Rust
cd src/rust
cargo build --release
cargo run --release
C#
cd src/csharp
dotnet build -c Release
dotnet run -c Release
F#
cd src/fsharp
dotnet build -c Release
dotnet run -c Release
Python (float-based)
cd src/python
python symbolic_bench.py
Python (NAND-recursive)
cd src/python
python nand_latin.py
Benchmarks
All implementations run the same mathematical pipeline on the same corpus (96 Latin verb forms, 23 Boolean features, 3 conjugations). Training: 3,000 epochs of full-batch gradient descent.
Speed
| Implementation | Training Time | Throughput | Inference Latency | Fwd+Bwd Latency |
|---|---|---|---|---|
| F# .NET 8 | 1.45s | 198,881 ex/s | 1.43 µs | 4.07 µs |
| C# .NET 8 | 2.25s | 127,879 ex/s | 2.02 µs | 6.49 µs |
| Rust (release) | 2.76s | 104,466 ex/s | 2.74 µs | 7.89 µs |
| Python (float) | 392s | 1,516 ex/s | 147.6 µs | 512.8 µs |
| Python (NAND) | 43.6s | ~37 ex/s | — | — |
xychart-beta
title "Training Throughput (examples/second)"
x-axis ["F#", "C#", "Rust", "Python\n(float)", "Python\n(NAND)"]
y-axis "examples/sec" 0 --> 210000
bar [198881, 127879, 104466, 1516, 37]
Cross-language run (identical data and initial weights)
The table above was measured on a different machine and architecture per runtime. To compare languages directly, bench/run.sh trains Rust, Mojo, BQN, Dyalog APL and Forth from the same bench/shared/corpus.txt and seed-42 bench/shared/init.txt: online SGD, lr 0.5, 3,001 epochs, embed 16 → tanh 32 → sigmoid 23. All of them finish at the same loss (9.1009e-5). Best of 3, single thread, x86-64 Linux.
| Rank | Runtime | Time | Throughput |
|---|---|---|---|
| 1 | Rust, allocation-free (src/rust/src/bin/lean.rs) |
0.62 s | 463,000 ex/s |
| 2 | Mojo 1.1.0 -O3 (mojo/morphology.mojo) |
0.91 s | 315,000 ex/s |
| 3 | Rust, original engine | 2.58 s | 112,000 ex/s |
| 4 | BQN, CBQN (bqn/morphology.bqn) |
2.77 s | 104,000 ex/s |
| 5 | Dyalog APL 20.0 (dyalog-apl/morphology.apls) |
5.91 s | 48,800 ex/s |
| 6 | Forth, gforth 0.7.3 (forth/morphology.fs) |
24.9 s | 11,600 ex/s |
xychart-beta
title "Cross-language training throughput (examples/second)"
x-axis ["Rust lean", "Mojo", "Rust orig", "BQN", "APL", "Forth"]
y-axis "examples/sec" 0 --> 500000
bar [463000, 315000, 112000, 104000, 48800, 11600]
Mojo beats the original Rust engine because that engine allocates in every forward and backward call; the allocation-free Rust port is the fair native baseline. The array languages pay per-primitive overhead on 16–32-wide vectors. Details in bench/results.md.
Quality
| Implementation | Train Accuracy | Unseen Accuracy | Notes |
|---|---|---|---|
| Rust | 96/96 (100%) | 83.2% feature-level | Mean-pool architecture |
| C# | 96/96 (100%) | 82.6% feature-level | Identical architecture to Rust |
| F# | 96/96 (100%) | 1/8 word-level | Fastest convergence |
| Python (float) | 196/198 (99%) | — | Elman RNN, 4 conjugations |
| Python (NAND) | 12/40 (30%) exact | 83.8% bit-level | Q(16.8) fixed-point via NAND |
Analysis
F# wins on speed. The .NET JIT optimizes F#'s expression-based, functional-style code more aggressively than C#'s mutable-statement style. F#'s immutable-by-default semantics give the optimizer more freedom to reorder, inline, and eliminate allocations.
C# and Rust are within 10% of each other. The .NET 8 JIT with AVX2 SIMD edges out LLVM's release build on this matrix-heavy workload. Both are ~85x faster than Python.
Python float is architecturally superior — it uses an Elman RNN that processes characters sequentially (crucial for suffix morphology) and trains on a richer corpus (198 words, 4 conjugations, passive voice, subjunctive). It pays ~140x for the interpreter.
The NAND variant decomposes every arithmetic operation to bit-parallel NAND gates through a Kogge-Stone carry-lookahead adder, shift-add multiplier, and table-lookup transcendentals in Q(16.8) fixed-point. It achieves 90% bit-level accuracy on seen words and 83.8% on unseen words with only D=3, H=4, and 40 epochs — a remarkable result for a system where every multiply is a software bit-serial loop.
Attention Kernel
Scaled dot-product attention, forward and backward. Rust: src/rust/src/attention.rs. Mojo: mojo/attention.mojo. F#: src/fsharp/Attention.fs. Python: src/python/attention.py. Full documentation: docs/ATTENTION.md.
Operation
Single head. Row-major f64 tensors, n rows × D columns. scale = 1/√D. With causal, entries with j > i are masked.
Forward Backward (given dO)
S = scale · Q Kᵀ dV = Pᵀ dO
P = softmax_row(S) dP = dO Vᵀ
O = P V Δᵢ = Σⱼ Pᵢⱼ dPᵢⱼ = dOᵢ · Oᵢ
Lᵢ = log Σⱼ exp(Sᵢⱼ) (saved) dS = P ⊙ (dP − Δ)
dQ = scale · dS K
dK = scale · dSᵀ Q
Implementations
| Path | Forward | Backward | Notes |
|---|---|---|---|
| Reference | materialised S, P | materialised P, dP, dS | n×n scratch; test oracle |
| Fused scalar | online softmax, 16-key blocks | two passes: dQ by query row, dK/dV by key row | P recomputed from saved L |
| AVX2+FMA | same | same | explicit intrinsics, runtime-detected |
| Threaded | row partition | row partition | bit-identical for any thread count |
lean.rs integration
lean [dir] [pool=attn|mean] [epochs=3001] [isa=auto|scalar|avx2] [threads=1]
pool |
Pooling step |
|---|---|
mean |
h = (1/n) Σᵢ emb[cᵢ] (original) |
attn (default) |
X = [emb[c₁]; …; emb[cₙ]], O = softmax(X Xᵀ/√16) X, h = (1/n) Σᵢ Oᵢ |
Backward for attn: dOᵢ = dh/n; gradient to letter position i = dQᵢ + dKᵢ + dVᵢ (Q = K = V = X).
Ports
| Language | Kernel | Tests | Paths |
|---|---|---|---|
| Rust | src/rust/src/attention.rs |
cargo test |
reference, scalar, AVX2+FMA, threaded |
| Mojo | mojo/attention.mojo |
mojo/attention_test.mojo |
reference, scalar, SIMD (4×f64 FMA) |
| F# | src/fsharp/Attention.fs |
selftest |
reference, scalar, AVX2+FMA (Vector256) |
| Python | src/python/attention.py |
python3 attention.py |
reference, scalar |
All four implement the same operation, the same xorshift64 input generator and the same 16-key online-softmax forward and two-pass backward, and are checked against bench/shared/attention_golden.txt.
Engine arguments
| Runtime | Invocation | Default pooling |
|---|---|---|
| Rust | lean [dir] [pool=attn|mean] [epochs=3001] [isa=auto|scalar|avx2] [threads=1] |
attn |
| Mojo | morphology [pool=mean|attn] [epochs=3001] [kernel=simd|scalar] |
mean |
| F# | dotnet run -c Release --project src/fsharp -- [mean|attn] [simd|scalar] |
mean |
Build, test, benchmark
cd src/rust
cargo test --release # kernel, lean.rs, CLI tests
cargo run --release --bin attention # correctness gate, then benchmark
cargo run --release --bin attention -- --write-golden
cargo run --release --bin lean -- ../../bench/shared attn
python3 ../python/attention.py # Python self-tests + golden vectors
cd mojo
mojo build attention_test.mojo -o /tmp/attention_test && /tmp/attention_test
mojo build -O3 attention_bench.mojo -o /tmp/attention_bench && /tmp/attention_bench
mojo build -O3 morphology.mojo -o /tmp/morph_mojo && /tmp/morph_mojo attn
dotnet run -c Release --project src/fsharp -- selftest # kernel + engine checks
dotnet run -c Release --project src/fsharp -- bench-attention
dotnet run -c Release --project src/fsharp -- attn
| Test | Suite |
|---|---|
| Fused vs reference, forward and backward (D ∈ {4, 8, 16, 64}, both masks, scalar and AVX2, 1–4 threads) | attention.rs |
| Central finite differences of dQ, dK, dV | attention.rs |
n = 0, n = 1, constant V, causal masking, large logits |
attention.rs |
| Bit-identical output for 1–256 threads | attention.rs |
| Golden vectors | attention.rs, attention.py, attention_test.mojo, selftest (F#) |
| Fused vs reference, finite differences, edge cases, error paths | attention_test.mojo, selftest (F#) |
| Engine embedding gradients vs finite differences (mean and attention pooling) | selftest (F#) |
Finite-difference check of embedding, w1, b2 gradients |
lean.rs |
pool=mean bit-identical to the original loop |
lean.rs |
| Scalar vs AVX2 training agreement | lean.rs |
| Argument handling, regression losses, determinism | lean_cli.rs |
Results
Median of 3 runs, 4-vCPU host. Single head, D = 64, n = 2048, causal.
| Implementation | Forward GF/s | Backward GF/s |
|---|---|---|
| Reference | 2.0 | 0.5 |
| Fused scalar, vectoriser off | 4.1 | 3.4 |
| Fused scalar, LLVM autovec | 4.8 | 4.4 |
| AVX2+FMA, 1 thread | 9.3 | 12.3 |
| AVX2+FMA, 4 threads | 25.8 | 30.9 |
Mojo and F#, D = 64, n = 2048, causal, median of 3 runs.
| Implementation | Forward GF/s | Backward GF/s |
|---|---|---|
| Mojo scalar | 3.0 | 2.5 |
| Mojo SIMD (4×f64 FMA) | 8.1 | 9.9 |
| F# scalar | 1.8 | 1.7 |
| F# AVX2+FMA | 5.3 | 6.1 |
lean, 96 words × 3001 epochs, best of 3.
| Pooling | Kernel | Time | Examples/s | Final loss |
|---|---|---|---|---|
| mean | — | 0.76 s | 381,000 | 9.1009e-5 |
| attention | scalar | 2.15 s | 134,000 | 8.3181e-5 |
| attention | AVX2+FMA | 1.40 s | 205,000 | 8.3181e-5 |
Mojo and F# training, 96 words × 3001 epochs, best of 3.
| Runtime | Pooling | Kernel | Time | Examples/s | Final loss |
|---|---|---|---|---|---|
| Mojo | mean | — | 1.14 s | 252,700 | 9.1009e-5 |
| Mojo | attention | SIMD | 2.76 s | 104,300 | 8.3181e-5 |
| Mojo | attention | scalar | 3.65 s | 78,900 | 8.3181e-5 |
| F# | mean | — | 2.41 s | 119,800 | 6.7e-5 |
| F# | attention | AVX2+FMA | 5.01 s | 57,500 | 6.5e-5 |
| F# | attention | scalar | 6.69 s | 43,100 | 6.5e-5 |
Mojo reads bench/shared/init.txt and corpus.txt, so its losses are directly comparable with Rust. F# initialises from Random(42).
The Three Layers
The architecture enforces a strict separation between observation, learning, and symbolic representation.
Observation Layer
The observation layer receives raw Latin text and converts it to numerical indices. No linguistic knowledge is encoded here — the characters are opaque symbols.
AMO → [65, 77, 79] (ASCII codes)
AMABAT → [65, 77, 65, 66, 65, 84]
REGUNT → [82, 69, 71, 85, 78, 84]
Variable-length inputs are handled naturally. The system makes no assumption about word length.
Learning Layer
The learning layer maintains:
Character embedding table: A 128×16 matrix (ASCII-indexed) mapping each character to a dense vector. These vectors are initialized randomly and learned from training data. Characters that appear in similar morphological contexts will develop similar embeddings.
Hidden layer weights: A 32×16 matrix W₁ and 32-dim bias b₁. The hidden layer computes nonlinear combinations of the pooled character representation.
Output layer weights: A 23×32 matrix W₂ and 23-dim bias b₂. The output layer maps hidden representations to grammatical feature predictions.
All parameters are updated by gradient descent after every training example.
Symbolic Layer
The symbolic layer converts continuous outputs to discrete Boolean values:
ŷᵢ ≥ 0.5 → TRUE (feature is present)
ŷᵢ < 0.5 → FALSE (feature is absent)
The threshold is fixed at 0.5. During training, we track the raw continuous values to measure convergence — how decisively the network commits to Boolean states.
Training Corpus
The corpus contains 96 word forms across 3 conjugation classes and 4 tense categories:
1st Conjugation (amāre, laudāre)
| Form | Person | Number | Tense |
|---|---|---|---|
| AMO / LAUDO | 1st | SG | PRESENT |
| AMAS / LAUDAS | 2nd | SG | PRESENT |
| AMAT / LAUDAT | 3rd | SG | PRESENT |
| AMAMUS / LAUDAMUS | 1st | PL | PRESENT |
| AMATIS / LAUDATIS | 2nd | PL | PRESENT |
| AMANT / LAUDANT | 3rd | PL | PRESENT |
| AMABAM / LAUDABAM | 1st | SG | IMPERFECT |
| ... | ... | ... | ... |
| AMAVI / LAUDAVI | 1st | SG | PERFECT |
| ... | ... | ... | ... |
2nd Conjugation (monēre, habēre)
Same paradigm structure with 2nd conjugation endings (-eo, -es, -et, -emus, -etis, -ent for present; -ebam, -ebas, -ebat for imperfect).
3rd Conjugation (regere, agere, ducere)
Same paradigm structure with 3rd conjugation endings (-o, -is, -it, -imus, -itis, -unt for present; -ebam, -ebas, -ebat for imperfect).
Generalization Test Set
8 words that do not appear in training:
| Word | Conjugation | Expected Features |
|---|---|---|
| NARRAT | 1st | 3sg present indicative active |
| NARRANT | 1st | 3pl present indicative active |
| NARRABAT | 1st | 3sg imperfect indicative active |
| VIDET | 2nd | 3sg present indicative active |
| VIDENT | 2nd | 3pl present indicative active |
| VIDEBAT | 2nd | 3sg imperfect indicative active |
| SCRIBIT | 3rd | 3sg present indicative active |
| SCRIBUNT | 3rd | 3pl present indicative active |
Boolean Convergence
During training, the network's outputs evolve from random (≈0.5 for all features) to decisive (≈0.0 or ≈1.0). We measure convergence by tracking the raw continuous values before thresholding.
Example: AMO after 3,000 epochs
PERSON:
PERSON_1 = 0.99982 → TRUE
PERSON_2 = 0.00001 → FALSE
PERSON_3 = 0.00004 → FALSE
NUMBER:
SINGULAR = 1.00000 → TRUE
PLURAL = 0.00000 → FALSE
TENSE:
PRESENT = 0.99993 → TRUE
IMPERFECT = 0.00000 → FALSE
FUTURE = 0.00015 → FALSE
PERFECT = 0.00003 → FALSE
MOOD:
INDICATIVE = 1.00000 → TRUE
SUBJUNCTIVE = 0.00001 → FALSE
IMPERATIVE = 0.00001 → FALSE
VOICE:
ACTIVE = 0.99999 → TRUE
PASSIVE = 0.00001 → FALSE
CONJUGATION:
CONJ_1 = 0.99946 → TRUE
CONJ_2 = 0.00067 → FALSE
CONJ_3 = 0.00001 → FALSE
All 23 features are correct. The raw values show decisive convergence — correct features at 0.999+, incorrect features at 0.001 or below. The network has learned that AMO is first person singular present indicative active first conjugation, not by memorization but by discovering that the character sequence A-M-O statistically correlates with these grammatical properties across the entire training set.
Generalization
The system's objective is not memorization. It must discover statistical relationships between character sequences and grammatical features that generalize to unseen words.
What the network learns
With mean-pool architecture, the network learns bag-of-character statistics — which characters tend to co-occur with which grammatical features. It discovers:
- The suffix
-muscorrelates with 1st person plural - The suffix
-ntcorrelates with 3rd person plural - The suffix
-bamcorrelates with imperfect tense - The suffix
-vi/-vitcorrelates with perfect tense - The vowel
ebefore endings correlates with 2nd conjugation - The absence of
abefore endings correlates with 3rd conjugation
Limitations
Mean pooling loses positional information. The network cannot distinguish AM (stem) from MA (reversed). An Elman RNN (as in the Python float variant) processes characters sequentially and can learn position-dependent suffix patterns more effectively.
The compiled implementations (Rust, C#, F#) use mean pooling for simplicity and speed. The Python float variant uses an Elman RNN and achieves higher accuracy on a richer corpus.
The NAND Variant
The file src/python/nand_latin.py implements the entire learning engine using a single primitive: the NAND gate.
NAND ← the only primitive
├── NOT, AND, OR, XOR, MUX
├── Kogge–Stone carry-lookahead adder
├── two's-complement negate / subtract
├── shift-add multiplier
├── Q(16.8) fixed-point representation
├── tanh / sigmoid (table + comparator + MUX)
└── the Latin morphology learner
├── character embedding
├── Elman RNN forward pass
├── binary cross-entropy loss
├── hand-written backpropagation through time
└── gradient descent θ ← θ − η ∇θ L
Every arithmetic operation — addition, multiplication, hyperbolic tangent, sigmoid — is built from bit-parallel NAND gates operating on 16-bit words. This is not a practical training engine (it runs at ~37 examples/second). It is a proof that the entire computation can be reduced to a single irreducible operation, the way silicon actually works.
Extending the System
Adding new verbs
Add entries to the corpus in dataset.rs / Dataset.cs / the F# Dataset module / the Python build_corpus() function. Each entry maps a word string to a target vector. The model architecture requires no changes.
Adding new features
Add feature names to the FEATURES array / Features.names / Features module. Increase the output dimension K. The network will learn to predict the new features alongside existing ones.
Changing the architecture
To switch from mean-pool to an Elman RNN (sequential processing):
- Replace the mean-pool step with a recurrent loop:
h_t = tanh(Wxh·x_t + Whh·h_{t-1} + bh) - Add BPTT (backpropagation through time) to the backward pass
- Add
Wxh,Whh,bhparameters
The Python float implementation already does this.
Increasing capacity
Increase EMBED_DIM (character embedding size) or HIDDEN_DIM (hidden layer size). More capacity allows the network to learn finer-grained patterns but requires more training data to avoid overfitting.
Mathematical Reference
Notation
| Symbol | Meaning |
|---|---|
| N | Word length (number of characters) |
| D | Embedding dimension (16) |
| H | Hidden dimension (32) |
| K | Output dimension (23) |
| E ∈ ℝ^{128×D} | Character embedding table |
| W₁ ∈ ℝ^{H×D} | Hidden layer weights |
| b₁ ∈ ℝ^H | Hidden layer bias |
| W₂ ∈ ℝ^{K×H} | Output layer weights |
| b₂ ∈ ℝ^K | Output bias |
| η | Learning rate (0.5) |
| σ(x) | Sigmoid: 1/(1+e^{-x}) |
| tanh(x) | Hyperbolic tangent |
Local Derivatives
| Operation | Forward | Local Derivative |
|---|---|---|
| Sigmoid | ŷ = σ(z) | dŷ/dz = ŷ(1-ŷ) |
| Tanh | a = tanh(z) | da/dz = 1 - a² |
| Linear | z = Wx + b | dz/dW = x, dz/dx = Wᵀ |
| Mean pool | h = (1/N)Σe_t | dh/de_t = 1/N |
| BCE loss | L = -[y log ŷ + (1-y) log(1-ŷ)] | dL/dŷ = -y/ŷ + (1-y)/(1-ŷ) |
Combined Derivatives
The sigmoid + BCE combination simplifies elegantly:
dL/dz₂ = (ŷ - y) / K
This is the gradient that flows backward from the output. It is clean because the sigmoid's derivative cancels with the BCE's denominator.
License
This project is licensed under the GNU Affero General Public License v3 (AGPL-3.0).
See LICENSE for the full text.
Every source file carries an AGPL-3.0 header. The AGPL requires that if you run a modified version of this software on a network server, you must make the source code available to users of that server. This ensures that improvements to the engine remain available to the community.
Symbolic Morphology Engine
Copyright (C) 2026 Ahmad Ali Parr
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU Affero General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
💼 Commercial License
SnapKitty code is free and open under AGPL-3.0 for open-source use. Building a commercial product or service? A proprietary commercial license from Snapkitty Collective LLC lets you ship this code without the AGPL's source-sharing and network-use obligations.
→ Get a commercial license · A.parr@belespritdaccord.uk
From letters to logic. From gradients to grammar. From NAND gates to noun declensions.