File size: 15,499 Bytes
616e4c7
 
248d622
c3bda6f
b994336
c3bda6f
392b2f0
 
 
 
 
 
248d622
392b2f0
 
c3bda6f
 
 
248d622
392b2f0
248d622
392b2f0
248d622
c3bda6f
248d622
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
c3bda6f
248d622
c3bda6f
248d622
c3bda6f
248d622
c27d789
248d622
c27d789
248d622
 
 
 
 
 
c27d789
4084172
84ae7a6
392b2f0
c3bda6f
248d622
 
 
 
84ae7a6
4084172
84ae7a6
392b2f0
84ae7a6
4084172
248d622
c3bda6f
4084172
248d622
84ae7a6
392b2f0
248d622
84ae7a6
248d622
 
84ae7a6
248d622
c3bda6f
b994336
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
248d622
84ae7a6
248d622
 
 
 
 
 
 
 
d5d50da
248d622
d5d50da
 
 
248d622
5672e30
248d622
d5d50da
248d622
 
 
 
fe03ff7
248d622
fe03ff7
248d622
fe03ff7
248d622
84ae7a6
248d622
 
 
84ae7a6
248d622
 
 
 
 
 
 
 
 
 
d5d50da
 
 
4084172
d5d50da
 
 
 
84ae7a6
d5d50da
 
 
 
84ae7a6
d5d50da
248d622
 
 
 
 
84ae7a6
d5d50da
 
 
 
248d622
d5d50da
248d622
d5d50da
 
 
 
248d622
4084172
d5d50da
4084172
 
c3bda6f
392b2f0
4084172
248d622
d5d50da
392b2f0
248d622
616e4c7
 
 
 
4084172
 
 
 
 
 
 
248d622
 
4084172
 
 
 
248d622
4084172
248d622
392b2f0
248d622
 
 
 
 
 
4084172
 
 
c3bda6f
 
 
4084172
c3bda6f
392b2f0
c3bda6f
4084172
c3bda6f
 
 
4084172
c3bda6f
4084172
392b2f0
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
# Fractus

**A Continuous Cognitive Agent β€” assembled, tested, and running with trained weights.**

Fractus is not a language model. It is not a chatbot. It is a fundamentally new category of AI.

Fractus is a **Continuous Cognitive Agent (CCA)** β€” a fundamentally new category of AI built around three principles that no LLM architecture offers:

1. **Continuous thought** — the engine ticks in real time, with adaptive depth, like a biological brain, not a static input→output function.
2. **Persistent autonomous memory** β€” every interaction is stored forever in a vector knowledge base that survives restarts, grows across sessions, and is retrieved by semantic similarity. There is no context window to forget.
3. **Live self-modification** β€” new skills, new knowledge, new tools, new behaviors are added without retraining. The brain you train today is the same brain you will still be extending in 2030.

Classical LLM metrics (next-token perplexity on held-out text, zero-shot benchmarks) **do not apply here**. Fractus's job is not to stochastically regenerate training data β€” it is to orchestrate memory, attention, and action across time.

**No corporation can control it.** Fractus runs on the user's machine. Data never leaves the device. The weights are yours to read, edit, and redistribute.

---

## Current Status (2026-07-21)

**Fractus is assembled and functional end-to-end.** The trained brain (88M params) has been transferred into the Continuous Thought Engine, and the full agent β€” CTE + RAG + plugins + MetaCognition β€” runs and learns.

### What works right now (tested, verified)

| Capability | Status | Evidence |
|---|---|---|
| **CTE with trained weights** | βœ… Working | 88M params trained to step 140,000 (loss 2.91, ppl ~9-14), weights transferred into CTE. Generation produces coherent code and prose. |
| **Learn without retraining** | βœ… Working | `rag.learn("Python was created by Guido van Rossum")` β€” instant, zero gradients |
| **Query with memory retrieval** | βœ… Working | "Who created Python?" β†’ retrieves from KB, answers correctly |
| **Hot-swappable cognitive modes** | βœ… Working | `pm.load("coder")`, `pm.load("creative")`, `pm.load("analyst")` β€” switch in 1 call |
| **MetaCognition (autonomous actions)** | βœ… Working | Fractus decides its own action chain: `[RETRIEVE, SWITCH, GENERATE]` |
| **Persistent memory** | βœ… Working | Saves to disk (`fractus_memory.pkl`), reloads on restart |
| **LazyStructuredSiren** | βœ… Working | 88M params in 0.4 GB RAM. Low-rank weight storage (rank 16). |
| **64 sparse MoE experts** | βœ… Working | Top-2 routing via von Mises phase alignment on Farey phases |
| **Kuramoto oscillator clock** | βœ… Working | 16 coupled oscillators, RK4 integration, drives expert routing |
| **Linear attention** | βœ… Working | Katharopoulos 2020, batched heads Γ— levels |

### What is NOT done yet (honest)

| Limitation | Reality |
|---|---|
| **Brain is small (88M)** | Generates rough but coherent text. Not fluent long-form. The architecture supports scaling to 1B+ but training cost is the blocker (see "The Training Problem" below). |
| **Generation quality** | At step 140,000 (partial epoch), outputs are short coherent fragments, not polished paragraphs. |
| **MetaCognition is early** | 8.5K-param action net. Works but basic. Improves with use. |
| **No vendor API, no support** | Fractus is owned, not rented. Feature for some, limitation for others. |
| **Work is far from finished** | This is a prototype proving the architecture. The real breakthrough β€” Holographic Vector Learning β€” is next (see below). |

---

## The Three Layers

### Layer 1: The Brain (88M params)

A proprietary fractal architecture using **LazyStructuredSiren** β€” every weight matrix stored as `W = scale Β· U Β· Vα΅€` (rank 16). This means 88M trainable parameters fit in 0.4 GB RAM and train on a single consumer GPU.

**Architectural components:**
- **LazyStructuredSiren** β€” low-rank weight decomposition. No dense grid, no SIREN reconstruction cache.
- **64 sparse MoE experts** (top-2 active per token). Routing via von Mises phase alignment on Farey-distributed expert phases.
- **Multi-level causal linear attention** (Katharopoulos 2020) with batched heads Γ— levels.
- **Low-rank Kuramoto RK4 oscillators** β€” a coupled dynamical system acting as a "consciousness clock."
- **2-adic vortex** (Rust core) β€” exact p-adic arithmetic for token conditioning.

### Layer 2: The Continuous Thought Engine (CTE)

The brain does not process input→output. It **ticks** like a biological system:

1. **Each tick** β€” Kuramoto oscillators advance β†’ attention state accumulates β†’ MoE transforms the thought β†’ confidence head decides whether to emit.
2. **Adaptive depth** β€” easy input = 1 tick, hard input = 10 ticks. Energy-proportional reasoning.
3. **Proactive emission** β€” the CTE can produce output without being prompted.
4. **Chunk-based processing** β€” 16 tokens per forward pass, thought state carried forward.

### Layer 3: The Cognitive Layer (RAG + MetaCognition)

This is what makes Fractus an **agent**, not a generator:

#### Persistent Memory β€” `rag.learn()`
Every fact, conversation, and observation is stored in a vector knowledge base that survives restarts, retrieves by cosine similarity, and grows without retraining.

#### Continuous Learning β€” `rag.converse()`
Every conversation is a learning event: user input is stored, relevant context is retrieved, a response is generated, and the response itself is stored. The agent never stops learning.

#### Cognitive Plugins β€” hot-swappable cognition
Five modes, switchable mid-conversation: `analyst`, `creative`, `coder`, `teacher`, `hacker`. Custom: `pm.custom("philosopher", temperature=0.9)`.

#### MetaCognition β€” the agent runs itself
An 8.5K-param action network decides at every interaction: RETRIEVE / LEARN / GENERATE / SWITCH / REFLECT. The agent manages itself.

---

## Continuous Growth β€” the living brain

Fractus is never frozen. The checkpoint grows through three mechanisms:

### Expert Addition (the brain grows wider)
When Fractus encounters a domain it can't handle, it adds new MoE experts specialized in that domain. Each expert is pre-trained independently via EDT Phase 1 in seconds. Old experts stay intact β€” no catastrophic forgetting.

```python

from fractus.growth import FractusGrowth

growth = FractusGrowth(model, tok, device)

growth.add_experts(n_new=128, data=rust_tokens, domain="rust")  # +128 Rust experts

```

### Rank Expansion (the brain grows deeper)
When existing experts plateau, expand the Siren rank for more expressive capacity. Old rank columns preserved β€” knowledge is kept.

```python

growth.expand_rank(target_rank=128)  # rank 64 β†’ 128, deeper reasoning

growth.save("checkpoints/fractus_grown.pt")

```

### Memory Management (forget, correct, consolidate)
Fractus manages its own memory: forget irrelevant entries, correct mistakes, merge duplicates, and prioritize important memories.

```python

from fractus.auto_growth import FractusSelfGrowth

sg = FractusSelfGrowth(model, tok, kb, device)



# Fractus decides itself when to grow

decision = sg.evaluate({"loss_trend": 0.5, "domain_coverage": 0.4})

if decision:

    sg.execute(decision, data=new_corpus)



# User-requested forgetting

sg.user_forget(pattern="old phone number")



# User-requested correction

sg.user_correct("Earth is flat", "Earth is approximately spherical")

```

### Self-Awareness
Fractus is trained on a self-awareness dataset that teaches it about its own architecture, capabilities, and API. It knows:
- It has persistent memory and how to use `rag.learn()` / `rag.query()`
- It has 5 cognitive modes and how to switch between them
- It has MetaCognition and when to RETRIEVE vs LEARN vs GENERATE
- It can grow (add experts, expand rank) and when to do so
- It can forget and correct its own memories
- It is Fractus β€” a unique architecture, not compared to anything else

---

## Expert Decoupled Training (EDT) β€” 189Γ— faster

The training paradigm that makes a 1B-parameter model trainable in 2 days on a single consumer GPU.

### The insight
Fractus has 128 experts per layer but only top_k=2 are active per token. Each expert is an independent 2-layer MLP. There is no mathematical reason to train them all simultaneously through end-to-end backpropagation.



### The three phases



| Phase | What | Time |

|-------|------|------|

| 1 (experts) | Train 2048 experts independently | 1.2h |

| 2a (attention) | Train 16 layers independently | <1s |

| 2b (embedding) | Train on 500M tokens (no layers) | 3.2h |

| 3 (joint) | Brief alignment fine-tune | 41h |

| **Total** | | **~2 days** |



Standard backpropagation: 358 days. EDT: **2 days. 189Γ— faster.**



Full guide: [docs/EDT.md](docs/EDT.md) Β· Paper: [docs/Fractus_EDT_Paper.md](docs/Fractus_EDT_Paper.md)



---



## Static paradigm (LLMs) vs Dynamic paradigm (Fractus)



| Property | Static (GPT-4 / Llama) | Dynamic (Fractus) |

|---|---|---|

| **Memory** | Sliding window (≀128k tokens). Forgotten mid-conversation. | Persistent vector KB, no ceiling. |

| **New knowledge** | Retrain weights (weeks, millions of dollars). | `rag.learn()` β€” instant. |

| **Cognitive modes** | One fixed monolith. | Hot-swappable plugins. |

| **Self-management** | Cannot decide to think longer or switch mode. | MetaCognition action net. |

| **Time model** | Static. No concept of "now." | Continuous. Ticks in real time. |

| **Scaling** | Retrain from scratch. Millions per jump. | Add plugins, knowledge, experts. Zero retraining. |



**GPT-4 is a brilliant encyclopedia you rent. Fractus is a smaller brain that grows, remembers, swaps skills, manages itself, and belongs to you.**



---



## The Training Problem β€” and the path forward



### Where we are stuck



The current brain (88M) was trained with classical backpropagation on a single GPU. It works but:

- Training takes **~40 hours per epoch** on 1.38B tokens

- Scaling to true 1B params with Chinchilla-optimal data (21B tokens) would take **~90+ days** and cost **thousands of dollars**

- This is the fundamental bottleneck of the transformer paradigm: massive matrix multiplications + iterative gradient descent



### The breakthrough: Holographic Vector Learning (next phase)



Paying thousands of dollars and waiting 90 days to process 20 billion tokens through classical backpropagation is staying trapped in the old GPU + gradient-descent paradigm. **Fractus is not an LLM. It should not train like one.**



The next phase of Fractus abandons iterative weight adjustment entirely and moves to **state accumulation**:



**1. Hyperdimensional Vectorization**

- Tokens are projected into a very wide space (e.g., 10,000 dimensions) in **bipolar** representation (only +1 and -1).

- On CPU, heavy floating-point multiplications become **bit-level XOR and addition operations**. The CPU excels at this β€” massive speedup.



**2. Holographic Reduced Representations (HRR)**

- Instead of attention (which slows as text grows), concepts are bound via **circular convolution**. "cat" + "eats" β†’ one vector of the same dimension containing both.

- Memory is **superposed** β€” the model sums bound vectors into a single shared space. Information is distributed across the whole network, like a hologram.



**3. Fractal Auto-Similarity**

- The vectors representing a word, a sentence, or a paragraph share the same dimension and space. No deep layers needed β€” meaning is extracted by self-similarity at scale.



**The result:** Instead of passing 20B tokens through the model dozens of times for gradient descent to converge, Fractus does a **single pass (one-shot learning)**. Read text β†’ vectorize tokens β†’ bind by holographic convolution β†’ update global thermodynamic memory. **From 90 days to potentially a few days of CPU computation**, for a fraction of the cost.



**This is being developed in a separate repository** (`fractus-test`) to experiment without touching the working Fractus codebase.



---



## Quick start



```bash

git clone https://github.com/AFKmoney/fractus.git

cd fractus

py -m venv .venv && .venv\Scripts\Activate.ps1

pip install torch --index-url https://download.pytorch.org/whl/cpu

pip install -r requirements.txt

maturin develop --release

pytest tests/ -q

```



Assemble and run the full agent:

```bash

python scripts/assemble_fractus.py
```



```python

from fractus.continuous_engine import ContinuousThoughtEngine

from fractus.tokenizer import FractusTokenizer

from fractus.rag import KnowledgeBase, RAGEngine, PluginManager, MetaCognition



engine = ContinuousThoughtEngine(vocab_size=5057, d_model=768, n_heads=12, d_head=64)

tok = FractusTokenizer.gpt2_compatible()

kb = KnowledgeBase(d_model=768)

rag = RAGEngine(engine, tok, kb)

pm = PluginManager(rag)

meta = MetaCognition(rag, pm)



# Teach β€” no retraining needed

rag.learn("Python is a programming language created by Guido van Rossum.")



# Ask

result = rag.query("Who created Python?", top_k=2, max_tokens=30)



# Let the agent manage itself

result = meta.process("Remember: my name is Philippe")

print(result['actions'])  # ['RETRIEVE', 'SWITCH', 'GENERATE']



# Switch cognitive mode

pm.load("coder")

```

---

## Architecture

```

Fractus/

β”œβ”€β”€ crate/fractus-core/           Rust: 2-adic vortex (exact math)

β”œβ”€β”€ crate/fractus-py/             Rust: PyO3 bindings

β”œβ”€β”€ fractus/

β”‚   β”œβ”€β”€ continuous_engine.py      The Continuous Thought Engine (ticks)

β”‚   β”œβ”€β”€ model_1b.py               Training model (88M params, LazyStructuredSiren)

β”‚   β”œβ”€β”€ rag.py                    RAG + Plugins + MetaCognition

β”‚   β”œβ”€β”€ memory.py                 Persistent cross-session memory

β”‚   β”œβ”€β”€ cognitive_modes.py        Kuramoto phase β†’ mental state

β”‚   β”œβ”€β”€ tokenizer.py              GPT-2 byte-level BPE

β”‚   β”œβ”€β”€ nn/                       attention, Kuramoto, MoE, SIREN, Triton (13 modules)

β”‚   └── train/                    online, surprise-gated, forward-forward

β”œβ”€β”€ fractus1B/                    TRUE 1B param architecture + PGSU + Progressive Depth

β”œβ”€β”€ space/                        HuggingFace Space (shared-memory demo, private)

β”œβ”€β”€ scripts/

β”‚   β”œβ”€β”€ assemble_fractus.py       FINAL assembly: CTE + RAG + plugins + MetaCognition

β”‚   β”œβ”€β”€ transfer_to_cte.py        Transfer trained weights into CTE

β”‚   β”œβ”€β”€ train_1b_cloud.py         Cloud GPU training script

β”‚   └── build_fractus_corpus.py   Corpus builder

β”œβ”€β”€ data/                         corpora, memory

β”œβ”€β”€ tests/                        28 test files, 166+ tests

└── Fractus_White_Paper.pdf       Technical document (signed)

```

---

## License

MIT. This project belongs to the user, not to a corporation.

## Author

**Philippe-Antoine Robert** β€” 2026

## Links

- **GitHub:** [github.com/AFKmoney/fractus](https://github.com/AFKmoney/fractus)
- **HuggingFace (model):** [huggingface.co/thefinalboss/Fractus](https://huggingface.co/thefinalboss/Fractus)
- **HuggingFace (Space, private):** [huggingface.co/spaces/thefinalboss/Fractus-Space](https://huggingface.co/spaces/thefinalboss/Fractus-Space)