| # Fractus | |
| **A Continuous Cognitive Agent β assembled, tested, and running with trained weights.** | |
| Fractus is not a language model. It is not a chatbot. It is a fundamentally new category of AI. | |
| Fractus is a **Continuous Cognitive Agent (CCA)** β a fundamentally new category of AI built around three principles that no LLM architecture offers: | |
| 1. **Continuous thought** β the engine ticks in real time, with adaptive depth, like a biological brain, not a static inputβoutput function. | |
| 2. **Persistent autonomous memory** β every interaction is stored forever in a vector knowledge base that survives restarts, grows across sessions, and is retrieved by semantic similarity. There is no context window to forget. | |
| 3. **Live self-modification** β new skills, new knowledge, new tools, new behaviors are added without retraining. The brain you train today is the same brain you will still be extending in 2030. | |
| Classical LLM metrics (next-token perplexity on held-out text, zero-shot benchmarks) **do not apply here**. Fractus's job is not to stochastically regenerate training data β it is to orchestrate memory, attention, and action across time. | |
| **No corporation can control it.** Fractus runs on the user's machine. Data never leaves the device. The weights are yours to read, edit, and redistribute. | |
| --- | |
| ## Current Status (2026-07-21) | |
| **Fractus is assembled and functional end-to-end.** The trained brain (88M params) has been transferred into the Continuous Thought Engine, and the full agent β CTE + RAG + plugins + MetaCognition β runs and learns. | |
| ### What works right now (tested, verified) | |
| | Capability | Status | Evidence | | |
| |---|---|---| | |
| | **CTE with trained weights** | β Working | 88M params trained to step 140,000 (loss 2.91, ppl ~9-14), weights transferred into CTE. Generation produces coherent code and prose. | | |
| | **Learn without retraining** | β Working | `rag.learn("Python was created by Guido van Rossum")` β instant, zero gradients | | |
| | **Query with memory retrieval** | β Working | "Who created Python?" β retrieves from KB, answers correctly | | |
| | **Hot-swappable cognitive modes** | β Working | `pm.load("coder")`, `pm.load("creative")`, `pm.load("analyst")` β switch in 1 call | | |
| | **MetaCognition (autonomous actions)** | β Working | Fractus decides its own action chain: `[RETRIEVE, SWITCH, GENERATE]` | | |
| | **Persistent memory** | β Working | Saves to disk (`fractus_memory.pkl`), reloads on restart | | |
| | **LazyStructuredSiren** | β Working | 88M params in 0.4 GB RAM. Low-rank weight storage (rank 16). | | |
| | **64 sparse MoE experts** | β Working | Top-2 routing via von Mises phase alignment on Farey phases | | |
| | **Kuramoto oscillator clock** | β Working | 16 coupled oscillators, RK4 integration, drives expert routing | | |
| | **Linear attention** | β Working | Katharopoulos 2020, batched heads Γ levels | | |
| ### What is NOT done yet (honest) | |
| | Limitation | Reality | | |
| |---|---| | |
| | **Brain is small (88M)** | Generates rough but coherent text. Not fluent long-form. The architecture supports scaling to 1B+ but training cost is the blocker (see "The Training Problem" below). | | |
| | **Generation quality** | At step 140,000 (partial epoch), outputs are short coherent fragments, not polished paragraphs. | | |
| | **MetaCognition is early** | 8.5K-param action net. Works but basic. Improves with use. | | |
| | **No vendor API, no support** | Fractus is owned, not rented. Feature for some, limitation for others. | | |
| | **Work is far from finished** | This is a prototype proving the architecture. The real breakthrough β Holographic Vector Learning β is next (see below). | | |
| --- | |
| ## The Three Layers | |
| ### Layer 1: The Brain (88M params) | |
| A proprietary fractal architecture using **LazyStructuredSiren** β every weight matrix stored as `W = scale Β· U Β· Vα΅` (rank 16). This means 88M trainable parameters fit in 0.4 GB RAM and train on a single consumer GPU. | |
| **Architectural components:** | |
| - **LazyStructuredSiren** β low-rank weight decomposition. No dense grid, no SIREN reconstruction cache. | |
| - **64 sparse MoE experts** (top-2 active per token). Routing via von Mises phase alignment on Farey-distributed expert phases. | |
| - **Multi-level causal linear attention** (Katharopoulos 2020) with batched heads Γ levels. | |
| - **Low-rank Kuramoto RK4 oscillators** β a coupled dynamical system acting as a "consciousness clock." | |
| - **2-adic vortex** (Rust core) β exact p-adic arithmetic for token conditioning. | |
| ### Layer 2: The Continuous Thought Engine (CTE) | |
| The brain does not process inputβoutput. It **ticks** like a biological system: | |
| 1. **Each tick** β Kuramoto oscillators advance β attention state accumulates β MoE transforms the thought β confidence head decides whether to emit. | |
| 2. **Adaptive depth** β easy input = 1 tick, hard input = 10 ticks. Energy-proportional reasoning. | |
| 3. **Proactive emission** β the CTE can produce output without being prompted. | |
| 4. **Chunk-based processing** β 16 tokens per forward pass, thought state carried forward. | |
| ### Layer 3: The Cognitive Layer (RAG + MetaCognition) | |
| This is what makes Fractus an **agent**, not a generator: | |
| #### Persistent Memory β `rag.learn()` | |
| Every fact, conversation, and observation is stored in a vector knowledge base that survives restarts, retrieves by cosine similarity, and grows without retraining. | |
| #### Continuous Learning β `rag.converse()` | |
| Every conversation is a learning event: user input is stored, relevant context is retrieved, a response is generated, and the response itself is stored. The agent never stops learning. | |
| #### Cognitive Plugins β hot-swappable cognition | |
| Five modes, switchable mid-conversation: `analyst`, `creative`, `coder`, `teacher`, `hacker`. Custom: `pm.custom("philosopher", temperature=0.9)`. | |
| #### MetaCognition β the agent runs itself | |
| An 8.5K-param action network decides at every interaction: RETRIEVE / LEARN / GENERATE / SWITCH / REFLECT. The agent manages itself. | |
| --- | |
| ## Continuous Growth β the living brain | |
| Fractus is never frozen. The checkpoint grows through three mechanisms: | |
| ### Expert Addition (the brain grows wider) | |
| When Fractus encounters a domain it can't handle, it adds new MoE experts specialized in that domain. Each expert is pre-trained independently via EDT Phase 1 in seconds. Old experts stay intact β no catastrophic forgetting. | |
| ```python | |
| from fractus.growth import FractusGrowth | |
| growth = FractusGrowth(model, tok, device) | |
| growth.add_experts(n_new=128, data=rust_tokens, domain="rust") # +128 Rust experts | |
| ``` | |
| ### Rank Expansion (the brain grows deeper) | |
| When existing experts plateau, expand the Siren rank for more expressive capacity. Old rank columns preserved β knowledge is kept. | |
| ```python | |
| growth.expand_rank(target_rank=128) # rank 64 β 128, deeper reasoning | |
| growth.save("checkpoints/fractus_grown.pt") | |
| ``` | |
| ### Memory Management (forget, correct, consolidate) | |
| Fractus manages its own memory: forget irrelevant entries, correct mistakes, merge duplicates, and prioritize important memories. | |
| ```python | |
| from fractus.auto_growth import FractusSelfGrowth | |
| sg = FractusSelfGrowth(model, tok, kb, device) | |
| # Fractus decides itself when to grow | |
| decision = sg.evaluate({"loss_trend": 0.5, "domain_coverage": 0.4}) | |
| if decision: | |
| sg.execute(decision, data=new_corpus) | |
| # User-requested forgetting | |
| sg.user_forget(pattern="old phone number") | |
| # User-requested correction | |
| sg.user_correct("Earth is flat", "Earth is approximately spherical") | |
| ``` | |
| ### Self-Awareness | |
| Fractus is trained on a self-awareness dataset that teaches it about its own architecture, capabilities, and API. It knows: | |
| - It has persistent memory and how to use `rag.learn()` / `rag.query()` | |
| - It has 5 cognitive modes and how to switch between them | |
| - It has MetaCognition and when to RETRIEVE vs LEARN vs GENERATE | |
| - It can grow (add experts, expand rank) and when to do so | |
| - It can forget and correct its own memories | |
| - It is Fractus β a unique architecture, not compared to anything else | |
| --- | |
| ## Expert Decoupled Training (EDT) β 189Γ faster | |
| The training paradigm that makes a 1B-parameter model trainable in 2 days on a single consumer GPU. | |
| ### The insight | |
| Fractus has 128 experts per layer but only top_k=2 are active per token. Each expert is an independent 2-layer MLP. There is no mathematical reason to train them all simultaneously through end-to-end backpropagation. | |
| ### The three phases | |
| | Phase | What | Time | | |
| |-------|------|------| | |
| | 1 (experts) | Train 2048 experts independently | 1.2h | | |
| | 2a (attention) | Train 16 layers independently | <1s | | |
| | 2b (embedding) | Train on 500M tokens (no layers) | 3.2h | | |
| | 3 (joint) | Brief alignment fine-tune | 41h | | |
| | **Total** | | **~2 days** | | |
| Standard backpropagation: 358 days. EDT: **2 days. 189Γ faster.** | |
| Full guide: [docs/EDT.md](docs/EDT.md) Β· Paper: [docs/Fractus_EDT_Paper.md](docs/Fractus_EDT_Paper.md) | |
| --- | |
| ## Static paradigm (LLMs) vs Dynamic paradigm (Fractus) | |
| | Property | Static (GPT-4 / Llama) | Dynamic (Fractus) | | |
| |---|---|---| | |
| | **Memory** | Sliding window (β€128k tokens). Forgotten mid-conversation. | Persistent vector KB, no ceiling. | | |
| | **New knowledge** | Retrain weights (weeks, millions of dollars). | `rag.learn()` β instant. | | |
| | **Cognitive modes** | One fixed monolith. | Hot-swappable plugins. | | |
| | **Self-management** | Cannot decide to think longer or switch mode. | MetaCognition action net. | | |
| | **Time model** | Static. No concept of "now." | Continuous. Ticks in real time. | | |
| | **Scaling** | Retrain from scratch. Millions per jump. | Add plugins, knowledge, experts. Zero retraining. | | |
| **GPT-4 is a brilliant encyclopedia you rent. Fractus is a smaller brain that grows, remembers, swaps skills, manages itself, and belongs to you.** | |
| --- | |
| ## The Training Problem β and the path forward | |
| ### Where we are stuck | |
| The current brain (88M) was trained with classical backpropagation on a single GPU. It works but: | |
| - Training takes **~40 hours per epoch** on 1.38B tokens | |
| - Scaling to true 1B params with Chinchilla-optimal data (21B tokens) would take **~90+ days** and cost **thousands of dollars** | |
| - This is the fundamental bottleneck of the transformer paradigm: massive matrix multiplications + iterative gradient descent | |
| ### The breakthrough: Holographic Vector Learning (next phase) | |
| Paying thousands of dollars and waiting 90 days to process 20 billion tokens through classical backpropagation is staying trapped in the old GPU + gradient-descent paradigm. **Fractus is not an LLM. It should not train like one.** | |
| The next phase of Fractus abandons iterative weight adjustment entirely and moves to **state accumulation**: | |
| **1. Hyperdimensional Vectorization** | |
| - Tokens are projected into a very wide space (e.g., 10,000 dimensions) in **bipolar** representation (only +1 and -1). | |
| - On CPU, heavy floating-point multiplications become **bit-level XOR and addition operations**. The CPU excels at this β massive speedup. | |
| **2. Holographic Reduced Representations (HRR)** | |
| - Instead of attention (which slows as text grows), concepts are bound via **circular convolution**. "cat" + "eats" β one vector of the same dimension containing both. | |
| - Memory is **superposed** β the model sums bound vectors into a single shared space. Information is distributed across the whole network, like a hologram. | |
| **3. Fractal Auto-Similarity** | |
| - The vectors representing a word, a sentence, or a paragraph share the same dimension and space. No deep layers needed β meaning is extracted by self-similarity at scale. | |
| **The result:** Instead of passing 20B tokens through the model dozens of times for gradient descent to converge, Fractus does a **single pass (one-shot learning)**. Read text β vectorize tokens β bind by holographic convolution β update global thermodynamic memory. **From 90 days to potentially a few days of CPU computation**, for a fraction of the cost. | |
| **This is being developed in a separate repository** (`fractus-test`) to experiment without touching the working Fractus codebase. | |
| --- | |
| ## Quick start | |
| ```bash | |
| git clone https://github.com/AFKmoney/fractus.git | |
| cd fractus | |
| py -m venv .venv && .venv\Scripts\Activate.ps1 | |
| pip install torch --index-url https://download.pytorch.org/whl/cpu | |
| pip install -r requirements.txt | |
| maturin develop --release | |
| pytest tests/ -q | |
| ``` | |
| Assemble and run the full agent: | |
| ```bash | |
| python scripts/assemble_fractus.py | |
| ``` | |
| ```python | |
| from fractus.continuous_engine import ContinuousThoughtEngine | |
| from fractus.tokenizer import FractusTokenizer | |
| from fractus.rag import KnowledgeBase, RAGEngine, PluginManager, MetaCognition | |
| engine = ContinuousThoughtEngine(vocab_size=5057, d_model=768, n_heads=12, d_head=64) | |
| tok = FractusTokenizer.gpt2_compatible() | |
| kb = KnowledgeBase(d_model=768) | |
| rag = RAGEngine(engine, tok, kb) | |
| pm = PluginManager(rag) | |
| meta = MetaCognition(rag, pm) | |
| # Teach β no retraining needed | |
| rag.learn("Python is a programming language created by Guido van Rossum.") | |
| # Ask | |
| result = rag.query("Who created Python?", top_k=2, max_tokens=30) | |
| # Let the agent manage itself | |
| result = meta.process("Remember: my name is Philippe") | |
| print(result['actions']) # ['RETRIEVE', 'SWITCH', 'GENERATE'] | |
| # Switch cognitive mode | |
| pm.load("coder") | |
| ``` | |
| --- | |
| ## Architecture | |
| ``` | |
| Fractus/ | |
| βββ crate/fractus-core/ Rust: 2-adic vortex (exact math) | |
| βββ crate/fractus-py/ Rust: PyO3 bindings | |
| βββ fractus/ | |
| β βββ continuous_engine.py The Continuous Thought Engine (ticks) | |
| β βββ model_1b.py Training model (88M params, LazyStructuredSiren) | |
| β βββ rag.py RAG + Plugins + MetaCognition | |
| β βββ memory.py Persistent cross-session memory | |
| β βββ cognitive_modes.py Kuramoto phase β mental state | |
| β βββ tokenizer.py GPT-2 byte-level BPE | |
| β βββ nn/ attention, Kuramoto, MoE, SIREN, Triton (13 modules) | |
| β βββ train/ online, surprise-gated, forward-forward | |
| βββ fractus1B/ TRUE 1B param architecture + PGSU + Progressive Depth | |
| βββ space/ HuggingFace Space (shared-memory demo, private) | |
| βββ scripts/ | |
| β βββ assemble_fractus.py FINAL assembly: CTE + RAG + plugins + MetaCognition | |
| β βββ transfer_to_cte.py Transfer trained weights into CTE | |
| β βββ train_1b_cloud.py Cloud GPU training script | |
| β βββ build_fractus_corpus.py Corpus builder | |
| βββ data/ corpora, memory | |
| βββ tests/ 28 test files, 166+ tests | |
| βββ Fractus_White_Paper.pdf Technical document (signed) | |
| ``` | |
| --- | |
| ## License | |
| MIT. This project belongs to the user, not to a corporation. | |
| ## Author | |
| **Philippe-Antoine Robert** β 2026 | |
| ## Links | |
| - **GitHub:** [github.com/AFKmoney/fractus](https://github.com/AFKmoney/fractus) | |
| - **HuggingFace (model):** [huggingface.co/thefinalboss/Fractus](https://huggingface.co/thefinalboss/Fractus) | |
| - **HuggingFace (Space, private):** [huggingface.co/spaces/thefinalboss/Fractus-Space](https://huggingface.co/spaces/thefinalboss/Fractus-Space) | |