Instructions to use nambor/refpred-operators with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nambor/refpred-operators with Transformers:
# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("nambor/refpred-operators", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 4,114 Bytes
279a795 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 | ---
license: other
tags:
- interpretability
- mechanistic-interpretability
- transformers
- residual-stream
---
# refpred-operators — fitted reference operators for transformer residual streams
Per-layer affine transition operators `A_l`, `b_l` fitted to the residual streams
of **Gemma-3-1b** and **Llama-3.2-1B**, by streaming ridge regression over a
corpus. They give you a **reference** for what each layer *usually* does to the
state, so you can score a single trajectory against it instead of against zero:
```
h_{l+1} ≈ A_l · h_l + b_l (the reference prediction)
ε_l = h_{l+1} − (A_l · h_l + b_l) (the innovation — what it does NOT predict)
```
This is the tool for the write-up *A Fitted Reference Predictor for Transformer
Residual-Stream Trajectories*. **You do not need to re-run the fitting pipeline** —
`torch.load` these and subtract from your own hidden states.
## Files
| file | contents |
|---|---|
| `operators_gemma.pt` | Gemma-3-1b, d=1152, 26 layers (~790 MB) |
| `operators_llama.pt` | Llama-3.2-1B, d=2048, 16 layers (~1.6 GB) |
| `load_operators.py` | tiny loader: `innovation(ops, h_l, h_next, layer)` |
Each `.pt` is a self-describing dict:
```python
{
"meta": { family, d, n_layers, hf_id_pt, hf_id_it, corpus, ridge_lambda, ... },
"PT-COMP": { "raw": {"A": [nL,d,d], "b": [nL,d]}, "normed": {...} },
"IT-COMP": { "raw": {...}, "normed": {...} },
"IT-CHAT": { "raw": {...}, "normed": {...} },
}
```
- **Conditions** — `PT-COMP` (base weights, no template), `IT-COMP` (instruct
weights, no template — the usual "deploy on the weights you'll use" default),
`IT-CHAT` (instruct weights, chat-templated).
- **Variants** — `raw` fits `hidden_states` as returned; `normed` fits
`input_layernorm_l(h_l)`, i.e. what block `l` actually reads. The two diverge
strongly on Gemma because of its `(1+γ)` RMSNorm — use `normed` if you feed
normalised states.
## Usage
```python
import torch
from load_operators import load_operators, innovation
ops = load_operators("operators_gemma.pt") # or operators_llama.pt
# your own hidden states for a prompt: hs[l] is [seq, d] (output_hidden_states=True)
layer = 12
eps = innovation(ops, hs[layer], hs[layer + 1], layer,
condition="IT-COMP", variant="raw") # [seq, d]
```
`predict_next(ops, h, layer, ...)` gives the reference prediction alone.
## Fit details
Ridge (λ_rel = 1e-6), fp32 extraction / fp64 solve, on fineweb-edu (sample-10BT,
seq_len 128, ~1400 fit docs). Operators are stored fp32. The BOS position is
excluded (attention sink); the chat scaffold is excluded by position for
`IT-CHAT`.
## What these are (and are not)
They are **regression coefficients fitted on activations** — not the model
weights, and not a runnable model. To produce hidden states you must load the
base model yourself (both are gated):
- Gemma-3-1b — `google/gemma-3-1b-pt` / `google/gemma-3-1b-it`,
**Gemma Terms of Use** (https://ai.google.dev/gemma/terms).
- Llama-3.2-1B — `meta-llama/Llama-3.2-1B` / `meta-llama/Llama-3.2-1B-Instruct`,
**Llama 3.2 Community License**.
Your use of these operators is subject to the respective base-model licenses.
## Caveats (read the write-up)
- The reference is **relative, not absolute**: an operator fitted on one domain
can be *worse than useless* on another (prose-vs-arithmetic differ ~58% on
shared support). Fit-distribution coverage is a real precondition.
- A large ε means "large," not "unusual" — turning it into an anomaly score needs
local calibration that these operators alone do not provide.
- Everything here is in the model's native basis; the geometry is
coordinate-dependent.
## Links & citation
- Write-up: *A Fitted Reference Predictor for Transformer Residual-Stream
Trajectories* — https://robman.fyi
- Code (fitting pipeline + analyses): see the accompanying repository.
```bibtex
@misc{manson_refpred_operators,
title = {Fitted reference operators for transformer residual streams},
author = {Manson, Rob},
year = {2026},
url = {https://robman.fyi}
}
```
|