content-address-v1

A pair of micro language models trained from scratch on a lookup task, demonstrating that content-addressable memory succeeds where attention-only retrieval fails.

Model Description

This model is a proof of concept for the working-memory bottleneck hypothesis described in Chapter 8 of Craft Erasure. The claim: an attention-only transformer must retrieve stored values by attending backward over the sequence, and this is bottlenecked by the effective attention window. A memory-augmented model that reads from a separate content-addressable store does not have this bottleneck, and its retrieval accuracy holds as the sequence grows.

Two models are trained on the same data with the same architecture and nearly identical parameter counts:

  • Model A (attention-only): standard causal transformer. Query must attend backward over the whole sequence to find the matching (position, value) pair.
  • Model B (memory-augmented): same transformer, plus a memory read embedding populated from a deterministic content-addressable store. At query positions, the read provides the value directly.

Task

[STORE p0 v0] [STORE p1 v1] ... [STORE pN vN] [QUERY pk]  ->  vk
  • p_i are position tokens 0-9
  • v_i are value tokens 0-9
  • pk is one of the positions seen earlier
  • The answer is the value token that was stored at pk

Each model is trained on sequences up to 6 stores and evaluated on sequences up to 16 stores. The key experiment is length generalization: does accuracy hold as the sequence grows beyond the training range?

Architecture

Component Value
Parameters (attention-only) ~25,000
Parameters (memory-augmented) ~25,500
Layers 2
Attention heads 2
d_model 32
d_ff 64
Memory dimension 32
Vocabulary 25 tokens
Max sequence length 64
Activation ReLU
Normalization LayerNorm (pre-norm)
Optimizer AdamW
Framework Pure numpy, no PyTorch

The memory-augmented model adds a single mem_emb matrix of shape (vocab_size, d_model). This is the only architectural difference. The PAD row of mem_emb is fixed at zero, so an empty memory read contributes nothing.

Memory Store

The content-addressable store used by Model B is deterministic and computed from the input:

  1. Walk through the sequence, tracking (position -> value) pairs as they appear after STORE tokens.
  2. When a QUERY token is followed by a position token, look up the stored value for that position.
  3. Emit the value token id as the memory read, or PAD if no read is available.

The model receives the memory read as an additional input feature at each position. It learns to interpret this feature during training. The store itself is not learned; only the embedding that maps stored values to vectors is trained.

Training Data

The training data is synthetic and generated on the fly. There is no external dataset. Each sequence contains 1-6 (position, value) pairs with distinct positions followed by a query for one of the stored positions. Values are random integers in [0, 9].

  • Training: 6,000 sequences, up to 6 stores each
  • Evaluation: 1,000 sequences with the same distribution (in-distribution check)
  • Length generalization: 200 sequences per length, lengths 1 to 16

Intended Uses

  • Research on memory architecture: study the difference between attention-based retrieval and content-addressable retrieval at minimal scale.
  • Length generalization studies: benchmark how well different architectures generalize beyond their training context length.
  • Education: a complete, self-contained demonstration of why working memory is a bottleneck and how a separate store removes it.
  • Starting point: extend to manipulation tasks (e.g., "value at position 5 plus 1") to test the working-memory slot explicitly.

How to Use

This model is not intended for production. It is a research and education artifact. To reproduce:

python content_address_v1.py

The script trains both models from scratch in ~5 minutes on a CPU and prints a comparison table showing the accuracy of each model at each sequence length.

There is no from_pretrained method. The models are defined and trained entirely within the script.

Expected Results

Sequence length Attention-only Memory-augmented
1-6 stores (in training range) ~100% ~100%
8-10 stores drops holds
12-16 stores near chance near 100%

The exact numbers depend on training seed and hyperparameters. The qualitative pattern is the claim: attention-only retrieval degrades with sequence length because the attention must span the growing working-memory window; memory-augmented retrieval holds because the store is content-addressable and has no window.

Limitations

  • Tiny scale. ~25k parameters. Not a language model. Not a general-purpose system.
  • Deterministic memory. The store is computed from the input, not learned. The model learns to use the store, not to build one. A full content-addressable memory model would need a learned write mechanism.
  • Single task. Retrieval of a stored value. No manipulation, no composition, no arithmetic on retrieved values.
  • No natural language. Inputs and outputs are synthetic token sequences.
  • Small positions. Only 10 distinct positions. Length generalization beyond 10 requires a larger position space.
  • No production safety. No alignment, no filtering, no refusal. Research artifact.

Why This Exists

Standard transformer architectures have a single working-memory workspace: the attention context window. Everything must pass through this window. For retrieval tasks, this means the query must attend backward over the whole sequence, and accuracy degrades as the sequence grows.

The claim from the paper is that content-addressable memory is a different cognitive operation โ€” direct access by address, not traversal of a sequence โ€” and that a model with an explicit content-addressable store can succeed where an attention-only model fails. This model is the smallest possible test of that claim.

The two models have nearly identical parameter counts and see identical data. The only difference is the presence of the memory read embedding and the store it draws from. If the memory-augmented model generalizes to longer sequences and the attention-only model does not, the working-memory bottleneck is demonstrated at laptop scale.

Citation

@software{content_address_v1,
  title = {content-address-v1: A Micro Language Model for Content-Addressable Memory},
  author = {[Author Names]},
  year = {2026},
  url = {https://huggingface.co/[user]/content-address-v1}
}

License

Apache 2.0. Weights, training script, and documentation are all released under the same license.

Acknowledgments

The architecture is motivated by the observation that human long-term memory is not bottlenecked by working memory in the way standard attention models are. This model is a minimal test of that observation. It is part of the broader project documented in Craft Erasure: How Fields Absorb Techniques and Discard the Minds That Produced Them.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support