content-address-v1
A pair of micro language models trained from scratch on a lookup task, demonstrating that content-addressable memory succeeds where attention-only retrieval fails.
Model Description
This model is a proof of concept for the working-memory bottleneck hypothesis described in Chapter 8 of Craft Erasure. The claim: an attention-only transformer must retrieve stored values by attending backward over the sequence, and this is bottlenecked by the effective attention window. A memory-augmented model that reads from a separate content-addressable store does not have this bottleneck, and its retrieval accuracy holds as the sequence grows.
Two models are trained on the same data with the same architecture and nearly identical parameter counts:
- Model A (attention-only): standard causal transformer. Query must attend backward over the whole sequence to find the matching
(position, value)pair. - Model B (memory-augmented): same transformer, plus a memory read embedding populated from a deterministic content-addressable store. At query positions, the read provides the value directly.
Task
[STORE p0 v0] [STORE p1 v1] ... [STORE pN vN] [QUERY pk] -> vk
p_iare position tokens0-9v_iare value tokens0-9pkis one of the positions seen earlier- The answer is the value token that was stored at
pk
Each model is trained on sequences up to 6 stores and evaluated on sequences up to 16 stores. The key experiment is length generalization: does accuracy hold as the sequence grows beyond the training range?
Architecture
| Component | Value |
|---|---|
| Parameters (attention-only) | ~25,000 |
| Parameters (memory-augmented) | ~25,500 |
| Layers | 2 |
| Attention heads | 2 |
| d_model | 32 |
| d_ff | 64 |
| Memory dimension | 32 |
| Vocabulary | 25 tokens |
| Max sequence length | 64 |
| Activation | ReLU |
| Normalization | LayerNorm (pre-norm) |
| Optimizer | AdamW |
| Framework | Pure numpy, no PyTorch |
The memory-augmented model adds a single mem_emb matrix of shape (vocab_size, d_model). This is the only architectural difference. The PAD row of mem_emb is fixed at zero, so an empty memory read contributes nothing.
Memory Store
The content-addressable store used by Model B is deterministic and computed from the input:
- Walk through the sequence, tracking
(position -> value)pairs as they appear afterSTOREtokens. - When a
QUERYtoken is followed by a position token, look up the stored value for that position. - Emit the value token id as the memory read, or
PADif no read is available.
The model receives the memory read as an additional input feature at each position. It learns to interpret this feature during training. The store itself is not learned; only the embedding that maps stored values to vectors is trained.
Training Data
The training data is synthetic and generated on the fly. There is no external dataset. Each sequence contains 1-6 (position, value) pairs with distinct positions followed by a query for one of the stored positions. Values are random integers in [0, 9].
- Training: 6,000 sequences, up to 6 stores each
- Evaluation: 1,000 sequences with the same distribution (in-distribution check)
- Length generalization: 200 sequences per length, lengths 1 to 16
Intended Uses
- Research on memory architecture: study the difference between attention-based retrieval and content-addressable retrieval at minimal scale.
- Length generalization studies: benchmark how well different architectures generalize beyond their training context length.
- Education: a complete, self-contained demonstration of why working memory is a bottleneck and how a separate store removes it.
- Starting point: extend to manipulation tasks (e.g., "value at position 5 plus 1") to test the working-memory slot explicitly.
How to Use
This model is not intended for production. It is a research and education artifact. To reproduce:
python content_address_v1.py
The script trains both models from scratch in ~5 minutes on a CPU and prints a comparison table showing the accuracy of each model at each sequence length.
There is no from_pretrained method. The models are defined and trained entirely within the script.
Expected Results
| Sequence length | Attention-only | Memory-augmented |
|---|---|---|
| 1-6 stores (in training range) | ~100% | ~100% |
| 8-10 stores | drops | holds |
| 12-16 stores | near chance | near 100% |
The exact numbers depend on training seed and hyperparameters. The qualitative pattern is the claim: attention-only retrieval degrades with sequence length because the attention must span the growing working-memory window; memory-augmented retrieval holds because the store is content-addressable and has no window.
Limitations
- Tiny scale. ~25k parameters. Not a language model. Not a general-purpose system.
- Deterministic memory. The store is computed from the input, not learned. The model learns to use the store, not to build one. A full content-addressable memory model would need a learned write mechanism.
- Single task. Retrieval of a stored value. No manipulation, no composition, no arithmetic on retrieved values.
- No natural language. Inputs and outputs are synthetic token sequences.
- Small positions. Only 10 distinct positions. Length generalization beyond 10 requires a larger position space.
- No production safety. No alignment, no filtering, no refusal. Research artifact.
Why This Exists
Standard transformer architectures have a single working-memory workspace: the attention context window. Everything must pass through this window. For retrieval tasks, this means the query must attend backward over the whole sequence, and accuracy degrades as the sequence grows.
The claim from the paper is that content-addressable memory is a different cognitive operation โ direct access by address, not traversal of a sequence โ and that a model with an explicit content-addressable store can succeed where an attention-only model fails. This model is the smallest possible test of that claim.
The two models have nearly identical parameter counts and see identical data. The only difference is the presence of the memory read embedding and the store it draws from. If the memory-augmented model generalizes to longer sequences and the attention-only model does not, the working-memory bottleneck is demonstrated at laptop scale.
Citation
@software{content_address_v1,
title = {content-address-v1: A Micro Language Model for Content-Addressable Memory},
author = {[Author Names]},
year = {2026},
url = {https://huggingface.co/[user]/content-address-v1}
}
License
Apache 2.0. Weights, training script, and documentation are all released under the same license.
Acknowledgments
The architecture is motivated by the observation that human long-term memory is not bottlenecked by working memory in the way standard attention models are. This model is a minimal test of that observation. It is part of the broader project documented in Craft Erasure: How Fields Absorb Techniques and Discard the Minds That Produced Them.