Planner-Cache / README.md
Dermitio's picture
Update README.md
e2d368d verified
|
Raw
History Blame Contribute Delete
7.61 kB
---
tags:
- semantic-memory
- planner-cache
- p-cache
- pytorch
- safetensors
- sqlite
- language-model-memory
- long-context
license: apache-2.0
---
# Planner Cache
Planner Cache is an external memory layer for frozen language models. It stores a bounded set of current facts outside the prompt, selects relevant facts, and exposes them through a small model or runtime compatibility artifact. This repository contains those Planner Cache artifacts. It is not a foundation model.
Planner Cache is not a foundation model and does not replace arbitrary long context. Recent KV, exact history, and tool retrieval remain responsibilities of the model runtime.
## Terms
- **P-cache** is a fixed-capacity store for facts that are currently true.
- **Canonical P** is the model-independent structured representation of those facts. It contains no model token IDs or hidden vectors.
- **Router or `.router`** selects the canonical state relevant to the current entity and relation.
- **TTL or `.ttl`** means Tensor Translation Layer. It converts canonical P into a model's internal state and represents semantic or internal support.
- **LTL or `.ltl`** means Lexical Translation Layer. It converts a routed value into tokenizer or output controls and represents lexical or output support.
- **P-package or `.ppkg`** stores durable personality patterns on disk and loads only selected entries.
- **Retained KV** is recent token-level attention memory maintained by the model runtime.
- An **active path** accepted memory and enabled TTL or LTL. An **inactive path** rejected, invalidated, or disabled memory and should match the frozen base.
- A **causal intervention** keeps the prompt fixed and changes only P. **KL divergence** measures how much the output distribution changed. **Incremental VRAM** is extra peak GPU memory above a warmed baseline.
![Planner Cache architecture](assets/architecture.svg)
## Distributed artifacts
| Artifact | Role | Compatibility |
|---|---|---|
| `canonical-p-v1.router` | Universal canonical state ranking and rejection | `pcm-canonical-p-v1` |
| `pythia-1.4b-final-layer.ttl` | Semantic query, value, gate, and layer metadata | Pythia-1.4B, hidden width 2048, layer 23 |
| `gemma4-e4b-q8-llama.ltl` | Direct adaptive lexical control metadata | Recorded Gemma4 Q8 GGUF and tokenizer checksums, llama.cpp |
| `personality-proof.ppkg` | Example durable personality package | `pcm-canonical-personality-v1` and canonical P v1 |
| Benchmark JSON | Raw recorded results | See [ARTIFACT_INDEX.md](ARTIFACT_INDEX.md) |
## Compatibility boundary
```text
canonical P state
-> universal .router
-> selected canonical value
-> native P, model-specific .ttl, or runtime-specific .ltl
-> frozen decoder
```
The canonical router has no model hidden dimension. A `.ttl` contains semantic compatibility weights. An `.ltl` contains lexical control metadata and may have no learned parameters. Neither contains base weights or P contents. The `.ppkg` contains canonical personality entries and evidence references but no model-native tensors.
The Gemma GGUF, tokenizer files, llama.cpp binaries, and upstream model metadata
are not redistributed. Users must supply the exact compatible bundle identified
by the LTL checksums.
## Loading the router and TTL
```python
from transformers import AutoModelForCausalLM
from pcm.planner import ByteEntityEncoder, CanonicalPRouter
from pcm.planner import PythiaSplitTranslatedModel, TensorTranslationLayer
base = AutoModelForCausalLM.from_pretrained(
"EleutherAI/pythia-1.4b",
dtype="float16",
).to("cuda")
ttl = TensorTranslationLayer.load(
"pythia-1.4b-final-layer.ttl",
device="cuda",
)
router = CanonicalPRouter.load("canonical-p-v1.router", device="cuda")
model = PythiaSplitTranslatedModel(base, ttl, router, ByteEntityEncoder())
```
The public Hub repository will need to provide the Planner Cache Python implementation or a pinned source release. The artifacts are not standalone Transformers models.
## Loading P-package
```python
from pcm.planner import PersonalityPackage, PersonalityQuery, PersonalityRouter
with PersonalityPackage("personality-proof.ppkg") as package:
selection = PersonalityRouter().retrieve(
package,
PersonalityQuery(
subject="user",
interaction_type="technical",
domain="debugging",
relation="response_style",
),
top_k=4,
)
```
## Measured results
The central causal result is that changing only valid P state changed the tested answer, while wrong, historical, invalidated, or disabled state left the tested base logits unchanged. At the 1,024-unit memory case, canonical P occupied about 2.05 MiB and retained KV tensors occupied about 193.31 MiB, a roughly 94-fold representation-size difference. These stores have different purposes and are not interchangeable.
- Post-audit canonical routing reached 100% top-1 and MRR 1.0 through 1,024 slots on the recorded synthetic scaling workload.
- The selected Pythia TTL reached 100% controlled held-out state generation at the 128-slot proof target.
- Wrong entity, wrong relation, historical, invalidated, router-disabled, and TTL-disabled matched CUDA conditions restored frozen-base logits where expected.
- The frozen Pythia base had zero parameters receiving gradients.
- The Gemma Q8 LTL emitted 128 of 128 held-out selected strings through direct adaptive logit bias. Inactive paths were exact.
- The Gemma LTL has zero learned parameters, occupies 708 bytes in the local artifact, adds no prompt tokens, and uses zero inactive VRAM.
- Indexed `.ppkg` header routing measured 67.10 ms at 100k entries and hydrated four rows from 152 headers.
- Opening an inactive `.ppkg` changed CUDA allocation by zero bytes.
- Natural post-turn review created and modified current RP state without explicit memory syntax in both interactive paths. It remains a controlled, slow extraction proof.
- The matched 1,024-token CUDA workload measured 103.491 MiB incremental peak for P-only with retained KV disabled, 294.266 MiB for retained KV only, and 294.783 MiB with both active.
See [TTL_CARD.md](TTL_CARD.md), [LTL_CARD.md](LTL_CARD.md), [ROUTER_CARD.md](ROUTER_CARD.md), [PPKG_CARD.md](PPKG_CARD.md), and the raw evidence in [ARTIFACT_INDEX.md](ARTIFACT_INDEX.md).
## Limitations
- Full trained semantic compatibility is proven only with Pythia-1.4B.
- The GPT-2 proof is structural and uses a tiny random model.
- Gemma has LTL support, not TTL support. Exact lexical emission does not establish internal semantic reasoning over P.
- A separate sequence-aware prototype reached 125 of 128 exact disjoint strings. It is not the selected runtime artifact and has not passed the 1,000-value or broad active-RP gates.
- The sequence complexity audit showed that exact performance comes primarily from tokenizer IDs and per-token forcing. Direct logit bias reached 128 of 128 with lower KL, but it is a lexical constraint rather than semantic translation.
- Canonical representation weights are reconstructed from a fixed recipe instead of being shipped as a standalone versioned artifact.
- Router-index hydration is linear in configured slot count.
- P-package personality behavior is a controlled deterministic proof, not nuanced neural personality learning.
- Planner Cache does not preserve exact old wording and does not replace archive retrieval.
- The reviewer currently focuses on owner, location, and status state. Review failures are inert and may miss valid facts. Recorded review latency was tens of seconds on the test hardware.