ensemble / README.md
thefinalboss's picture
v0.5.1: morphological filtering in solve_analogy (50%->62%), cosmul option exposed, cosadd default
593be8a verified
|
Raw
History Blame Contribute Delete
15.1 kB
---
license: mit
language: en
tags:
- hyperdimensional-computing
- vector-symbolic-architectures
- associative-memory
- training-free
- kuramoto
- ensemble-methods
- cpu
- no-gpu
- memory-augmented
- bpe
library_name: ensemble
---
# ENSEMBLE — a training-free AI built from portable experts
> **No gradient. No epochs. No GPU.** Turn any dataset into a compressed expert file in one pass, then let a Kuramoto-coupled brain synchronize experts into emergent answers. ENSEMBLE thinks when idle, **grows a persistent central memory**, and safely rewrites its own coupling. It is built on the [Palimpseste](https://huggingface.co/thefinalboss/palimpseste-max) hypervector substrate.
ENSEMBLE is **not** a transformer and is not trying to be one. It is an experiment in a different direction: a training-free, ultra-compact, instantly-updateable, composable associative memory that grows smarter over time without retraining. This README is precise about what it is, what it does well, and where it loses to a trained 1B transformer.
## What's new in v0.2
- **Global shared BPE tokenizer** — one tokenizer across all experts lifts the char-level quality ceiling. On TinyStories (150 KB excerpt, D=5000): **next-token accuracy 94.5% → 98.5%**, **query latency 112ms → 28ms (4× faster)**, **RAM 187MB → 89MB (2× less)**.
- **Persistent central brain memory** — the brain now accumulates synthetic concepts (from continuous thinking) in a memory that survives `save()` / `load()`. The system grows indefinitely: add experts + let it think, no retraining.
- **Brain save/load** — full brain state persistence.
- 75 tests passing (was 49).
## What's new in v0.3 — structural generalization
- **Structural query encoding** — decompose a question into a PATTERN + a SLOT via explicit templates (`Expert.from_qa_pairs(..., patterns=["what is the capital of {country}"])`). Questions sharing a pattern are encoded consistently, so the answer **format generalizes to unseen slots**.
- **Measured generalization lift**: on capitals (train: france/germany/italy/japan/egypt; holdout: spain/portugal/greece/brazil/norway/india/mexico/kenya), the char-level baseline returns an **empty answer 100% of the time** on unseen slots; the structural encoder returns a **non-empty answer 100% of the time**. That is **+100 points** of graceful degradation on the unknown.
- **Honest scope**: this guarantees a well-formed answer shape, *not* factual correctness for a fully unseen slot. Without semantic embeddings, "spain" and "france" are orthogonal at the HV level, so the system imitates structure rather than interpolating meaning. The README is explicit about this.
- 91 tests passing (was 75).
## What's new in v0.4 — embedding expert (optional semantics)
- **Embedding expert** — an optional plugin that brings real **semantic similarity** to the brain. Dense word vectors (fastText, or a local mini-embedding) are projected into HV space via a seeded random-projection-then-sign (the new `projection.py` primitive), so `spain` and `portugal` become **similar hypervectors** (cosine-preserving, JL-style). This is exactly the ingredient the structural encoder needed for analogy to *find the right slot*.
- **It's a pure plugin**: `EmbeddingExpert` satisfies the same duck-typed contract as `Expert` (relevance / candidate_hv / answer / signature_hv), so it plugs into the Brain and couples with lexical experts via Kuramoto **with zero changes to the brain**. The training-free / numpy-only core stays pure; semantics is opt-in.
- **Two construction paths**: `from_corpus_local` (PPMI+SVD mini-embedding, training-free, numpy only, seconds — lower quality) and `from_fasttext` (load pretrained `.vec`, best quality, one-time download).
- **Honest result**: the embedding expert measurably brings semantic similarity (`spain~portugal = 0.18` vs `~0` without), which improves routing. However, with the **mini-embedding** the factual-correctness on the capitals holdout remains ~0% — the slot is found but the answer content is still decoded char-level. The pipeline is ready for fastText to deliver real factual gains; that's the documented next step.
- 106 tests passing (was 91).
## What's new in v0.5 — factual generalization via dense analogy
- **Dense analogy solving** — `EmbeddingExpert.learn_relation("capital_of", {france: paris, ...})` + `solve_analogy("capital_of", "spain")` implements the classic Mikolov analogy (`paris − france + spain ≈ madrid`) in normalized dense space, averaging the `(answer − slot)` direction over known examples and finding the nearest word — **excluding the query slot** so it doesn't parrot itself.
- **Wired into structural answers** — when an embedding expert is attached to a lexical expert (`from_qa_pairs(patterns=..., embedding=emb)`), unseen slots are answered via dense analogy instead of char-level decode, returning the **plain-text** answer.
- **Measured factual gains (fastText, end-to-end via the Brain)**: on a capitals holdout (train: france/germany/italy/japan/egypt/russia; holdout: spain/portugal/greece/norway/china/india/brazil/turkey), factual correctness goes from **0% (char) / 0% (structural-only) to ~50–67% (structural + embedding analogy)**. Concretely: `spain → madrid`, `portugal → lisbon`, `greece → athens`, `norway → oslo` are recovered correctly despite never being in training.
- **Morphological filtering** (default on): excludes candidates sharing the query slot's stem (filters `indian`, `turkeys`, `brazilian`), which empirically recovers `india → delhi` (50% → 62% on the isolated analogy benchmark). Two scoring methods exposed: `cosadd` (default, Mikolov 3CosAdd) and `cosmul` (Levy & Goldberg 3CosMul) — measured `cosadd+morph` to be the best on this benchmark.
- **Honest scope**: fastText analogies are imperfect (china → porcelain on a miss) — that's fastText's distributional ceiling, not the architecture. More known pairs and larger vocab would lift it further. The point stands: the wall is broken.
- 109 tests passing (was 106).
---
## The four ideas
### 1. Dataset → Expert (`.exp`) — one pass, no gradient, smaller than the source
Any dataset becomes a frozen, portable expert file in a single pass. There is no gradient descent and no epoch — learning is an `O(1)`-per-token write into an append-only memory. The expert file is **smaller than the dataset** (typically 3–9×): hypervectors are never stored; the symbolic token stream is gzipped and HVs are rebuilt on load from a deterministic encoder.
### 2. Kuramoto brain — experts couple and synchronize into emergent answers
Load several experts into a **Brain**. Each expert becomes a Kuramoto oscillator (natural frequency = relevance, coupling = signature similarity). The oscillators synchronize by similarity; the emergent attractor is a state **no single expert produced**. The brain routes questions to the right expert or composes answers when two experts know the topic.
### 3. Global shared BPE + persistent central memory — the brain grows
A single BPE tokenizer makes all experts token-compatible (so they share the same vocabulary and can be assembled freely). Meanwhile the brain's **central memory** accumulates synthetic concepts discovered during continuous thinking — and unlike a transformer's frozen weights, this memory **persists and grows**. Adding knowledge is Lego: drop in a `.exp`, let the brain think, and its concept store expands. No joint retraining.
### 4. Continuous thought + safe self-modification
When idle, the brain **thinks**: it samples remembered queries from one expert, asks all the others, runs the attractor, and writes novel coherent results into the central memory. It can also **rewrite its own coupling** under a **Lyapunov constraint** (ΔE ≤ 0). The acceptance rule is immutable — the recursion is bounded by construction.
---
## How it differs from an LLM
| | Transformer LLM | ENSEMBLE |
|---|---|---|
| Learning | gradient descent, ~1T tokens | one-pass write, no gradient |
| Adding knowledge | retrain / fine-tune / RAG | drop in a `.exp` (Lego) + think |
| Model size vs data | grows with parameters | expert is **smaller** than its data |
| Grows over time | frozen weights | **central memory accumulates concepts** |
| Idle behavior | nothing | dreams across experts, writes concepts |
| Self-tuning | hyperparameters fixed | rewrites its own coupling (Lyapunov-bound) |
| Hardware | GPU | plain CPU |
---
## Quick start
This repo is **self-contained**: it vendors the `palimseste` substrate, so a fresh clone runs with only `numpy`.
```bash
git clone https://huggingface.co/thefinalboss/ensemble
cd ensemble
pip install numpy # the only runtime dependency
python -c "from ensemble import Brain, Expert; print('ok')"
```
### Build experts and assemble a brain
```python
from ensemble import Expert, Brain
# Option A: char-level (backward compatible, simplest)
math = Expert.from_qa_pairs(
[("what is pi", "pi is approximately three point one four")] * 5,
domain="math", D=10000)
# Option B: global shared BPE (recommended — higher quality)
corpus = open("some_corpus.txt").read()
bpe = Expert.build_bpe(corpus, vocab_size=2000, D=10000) # train once
math = Expert.from_qa_pairs(
[("what is pi", "pi is approximately three point one four")] * 5,
domain="math", D=10000, tokenizer=bpe) # share it
geo = Expert.from_qa_pairs(
[("what is the capital of france", "the capital of france is paris")] * 5,
domain="geography", D=10000, tokenizer=bpe)
brain = Brain()
brain.add_expert(math)
brain.add_expert(geo)
print(brain.query("what is pi").answer) # -> math
print(brain.query("what is the capital of france").answer) # -> geography
brain.think(seconds=10) # the brain dreams -> writes to central memory
brain.self_modify() # safely retunes its own coupling
# the central memory persists
brain.save("mybrain")
brain2 = Brain.load("mybrain") # concepts survive reload
print(f"{brain2.n_concepts} concepts persisted")
```
### Saving experts (compressed)
```python
result = math.save("math.exp") # source -> expert, compressed
math2 = Expert.load("math.exp", tokenizer=bpe) # BPE experts need the shared tokenizer
```
### Structural generalization (unseen slots)
```python
# Train on capitals of some countries, WITH a template.
# The template lets the expert generalize the *answer format* to unseen slots.
geo = Expert.from_qa_pairs(
[("what is the capital of france", "the capital of france is paris"),
("what is the capital of germany", "the capital of germany is berlin")] * 4,
domain="geo", D=10000,
patterns=["what is the capital of {country}"])
# spain was NEVER in training — yet the structural expert answers (format-generalized),
# where a char-level expert would return empty.
print(geo.answer("what is the capital of spain")) # non-empty, by analogy
```
---
## Scaling to "1B-equivalent"
ENSEMBLE has **no stored parameters**, so "1B" means **capacity** — distinct associations the memory holds without collision, exponential in dimension `D`. The `1b` preset uses `D = 100 000`.
| preset | D | use case |
|---|---|---|
| `tiny` | 2,000 | quick demos |
| `small` | 10,000 | laptop default |
| `medium` | 30,000 | more capacity |
| `large` | 50,000 | large corpora |
| `1b` | 100,000 | "1B-equivalent" capacity |
---
## Benchmark results (measured, reproducible)
Full methodology in [`RESULTS.md`](./RESULTS.md).
### BPE vs char-level on TinyStories (150 KB excerpt, D=5000)
| metric | char | **BPE** | delta |
|---|--:|--:|---|
| **next-token accuracy** | 94.5% | **98.5%** | **+4.0 pts** |
| **query latency** | 112 ms | **28 ms** | **4× faster** |
| **RAM** | 187 MB | **89 MB** | **2× less** |
| build time | 46s | 107s | one-shot cost |
The global BPE lifts the quality ceiling exactly as predicted: higher accuracy, far lower latency (fewer, longer tokens), and half the RAM. The tradeoff is a slower one-time build.
### ENSEMBLE scaling with D (facts corpus)
| D | next-token acc | QA recall | .exp/source |
|--:|--:|--:|--:|
| 2,000 | 30.6% | 0% | 7.0× |
| 10,000 | 93.5% | 35% | 7.0× |
| 100,000 (1b) | 94.7% | 40% | 7.0× |
### Generalization to unseen slots (capitals benchmark)
Train on capitals of 5 countries; hold out 8 unseen countries. Metric: % of holdout returning a non-empty, well-formed answer.
| mode | non-empty on unseen slots |
|---|--:|
| char-level | **0%** (silent failure) |
| **structural (patterns)** | **100%** (graceful, format-generalized) |
Honest: the structural answers are well-formed guesses by analogy, not factually correct for fully unseen slots (no semantic embeddings). The win is *graceful degradation* — the system says something sensible instead of nothing.
### Honest verdict vs 1B transformers (TinyLlama-1.1B, Pythia-1B)
**ENSEMBLE wins on:** zero training (~30s CPU vs a GPU cluster on 1T tokens), footprint (a `1b` expert is **~2 KB** vs **~2 GB**), instant knowledge injection, compositionality, near-perfect memorization of seen data (95–99% with BPE), and **continuous growth** (the central memory never stops accumulating).
**A 1B transformer wins on:** broad world knowledge (MMLU, HellaSwag), **generalization** (ENSEMBLE memorizes; holdout QA ≈ 0%), fluency on unseen text, reasoning.
**Bottom line:** ENSEMBLE is a different tool. For narrow domains with known data and a CPU-only / tiny-footprint constraint, it is competitive or superior. For general intelligence, it is not — yet.
---
## Architecture
```
Expert dataset -> compressed .exp (one pass, no gradient)
â”” build_bpe train a global shared BPE tokenizer
ExpertOscillator an expert as a Kuramoto oscillator
ExpertKuramotoAttractor multi-expert synchronization -> emergent attractor
Brain conductor: roster, query, think, self-modify, save/load
BrainMemory persistent central memory (grows via thinking)
ContinuousThinking cross-expert dreaming -> writes concepts to BrainMemory
SelfModifier Lyapunov-bounded retuning of coupling
```
---
## What this is (and is not)
**Is:** a training-free associative memory built on hypervector (VSA) algebra; an expert system where knowledge is portable, compressed, composable, and grows over time; a research artifact exploring Kuramoto coupling + persistent concept memory. CPU-only, numpy-only, auditable.
**Is not:** a transformer; competitive on broad-knowledge benchmarks; a generalizer (it memorizes what it is shown).
---
## Run the tests
```bash
pip install numpy pytest
pytest tests/ -q # 75 tests
```
## Repository layout
```
ensemble/ the ENSEMBLE package (expert, bpe integration, brain, brain_memory, ...)
palimseste/ the vendored hypervector substrate
bench/ reproducible benchmarks (incl. BPE vs char comparison)
tests/ 75 tests
RESULTS.md full benchmark report
```
## License
MIT. Both `ensemble` and the vendored `palimseste` substrate are MIT-licensed.