maxact-fast / README.md
ceselder's picture
SPEED.md: MFU state + open speed levers for hill-climbing; make full README canonical
8e0dddc
|
Raw
History Blame Contribute Delete
2.74 kB
---
license: apache-2.0
tags:
- interpretability
- sae
- activation-steering
- max-activating-examples
---
# maxact-fast: a general residual-direction β†’ text inverter
**It's all just probes.** SAE encoder columns, cluster classifiers, hand-crafted concept vectors β€”
every one of them is a unit direction in a language model's residual stream. This project trains
Qwen3-8B (LoRA) to take *any* such direction β€” injected activation-oracle-style into its own
layer-1 residual stream β€” and **write the text that maximally activates it** at layer 27.
Prior work (our SAE line) showed a generator trained on one SAE's features beats exhaustive
500M-token corpus search on ~30% of held-out features at best-of-16, and β€” via a Sonnet-judged
audit of all 65k features β€” that corpus max-activating examples alone *mislabel or under-specify
~72% of features* (56.8% refined, 3.6% outright contradicted, with 918 contradictions surviving
adversarial re-testing). The bottleneck is conditioning diversity: one SAE = 65k directions.
## The direction firehose
1. **Embed** UltraFineWeb (streamed) with BGE β†’ **k-means K clusters** (pilot K=10k, target K=1M).
2. Sample cluster pairs (A,B), **train a logistic-regression probe** on the model's own layer-27
residuals; keep the direction only if it *validates* (held-out AUC β‰₯ 0.9 and cluster A actually
projects high on it β€” no junk directions).
3. **SFT target** = corpus texts nearest A's centroid: "given direction w, write text like this."
4. **Pretrain** on millions of validated (direction, text) pairs β†’ **post-train** on SAE-feature
max-activating corpus examples β†’ **RL**.
## RL: Dr. GRPO, no KL
Advantage = `r βˆ’ group_mean` (no /std β€” the Dr. GRPO fix), loss token-summed over a global batch
normalizer, importance ratios TIS-capped, **no KL leash** (it blocked exactly the rare-feature
specialization we want) β€” diversity comes from an explicit entropy bonus (`maximize r + Ξ²Β·H`)
plus fluency/distinct gates. Rollouts via vLLM (`vllm-lens` per-request steering injection).
Reward = max-over-positions dot of the *standalone re-tokenized* generation through the clean
base model. Fair by construction.
## Why
If the cross-uplift bet lands β€” cluster-probe pretraining transferring zero-shot to SAE features β€”
you get a **universal activation-maximization oracle**: hand it any probe you trained, get back
legible text showing what that probe actually fires on. Feature auditing, probe debugging, and
"what did my classifier really learn" become one generate() call.
Code in this repo; SAE-line results: [qwen3-8b-sae-maxact-lora](https://huggingface.co/ceselder/qwen3-8b-sae-maxact-lora).
MATS interpretability project (Neel Nanda stream).