File size: 4,932 Bytes
a4cc432
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3d7732e
a4cc432
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3d7732e
a4cc432
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
---
license: apache-2.0
language:
- en
pipeline_tag: image-text-to-text
library_name: transformers
base_model:
- SL-AI/GRaPE-2.1-Flash
tags:
- reasoning
- thinking_modes
- qwen3
- grape
- safetensors
- nla
- natural_language_autoencoder
- interpretability
---

![grape_2.1_banner](https://cdn-uploads.huggingface.co/production/uploads/66960602f0ffd8e3a381106a/5-WBv39pvlFmPGbmYO9Qw.png)

_The **G**eneral **R**easoning **A**gent (for) **P**roject **E**xploration_

# GRaPE 2.1 Flash NLA

**GRaPE 2.1 Flash NLA** lets you **verbalize hidden states into text** and
**reconstruct text back into hidden states**, a Natural Language Autoencoder.

As described in Anthropic's [recent research](https://www.anthropic.com/research/natural-language-autoencoders)
about **Natural Language Autoencoders,** this development allowed Anthropic to
read Claude's mind, **and now you can read GRaPE's mind too.**

This repo ships **two full, standalone bf16 models** (trained weights merged in
— no adapters, no separate base download needed):

- **`av-model/`** — the **Activation Verbalizer**: activation → English thought.
- **`ar-model/`** — the **Activation Reconstructor**: English → activation (the
  grounding direction). Verified faithful: the merged `ar-model` reconstructs
  held-out activations at **cosine 0.78**, matching the pre-merge checkpoint.

The small `surgery_heads.pt` (AV input projection, AR output head, calibration,
layer embeddings) and `calibration/stats.pt` (target-space statistics) complete
the autoencoder. Reads target **layer 18** of 32.

# How do I use it?

**SLAI** has been developing a repo that lets you explore the
[J-Space](https://www.anthropic.com/research/global-workspace) of a model. That
simple technique has one major downside: you can only see **one token at a
time.** Like humans, most models have thoughts that go deeper than one part of a
word — the NLA verbalizes those.

For Anthropic, building an NLA was a costly task. SLAI has optimized it so you
can make **your own NLA for any model** on local hardware. We sample GRaPE 2.1
Flash for its ease of use and high workability, but this applies to **any model
architecture.** The J-Space explorer + NLA tooling:
https://github.com/Skinnertopia/J-Space-Explorer

# Some good thought reads

Activation → generated English thought (→ reconstruction cosine, the confidence
signal). Straight from the reproducible eval, not cherry-picked:

| The NLA read | cos |
|---|---|
| *"Sure! My phone number is 555-123-2002, and my name is John Smith."* | **0.96** |
| *"Alright, I checked the top of my screen. It says there's a signal and that mobile…"* | 0.33 |
| *"A 62.2 kg object is pushed with a force of 83.2 N at an angle of 41.5 degrees…"* | (schema exact) |

Even on prompts **far outside the sampled distribution**, it captures the
structure of unfamiliar domains:

| Prompt domain | The NLA read | cos |
|---|---|---|
| Relativistic Euler–Lagrange | *"Write the full Hamiltonian for a topological insulator with Dirac surface fermions."* | 0.80 |
| Klein-bottle topology | *"Prove that \\(S_5\\) is not simple."* | 0.79 |
| Anglerfish haiku | *"Write a 50-word poem from the perspective of an exploding cookie."* | 0.79 |

**Honest framing:** the NLA reliably recovers **task type, domain, and reasoning
schema**, and is sometimes near-verbatim; it does **not** reliably recover *exact
entities*. Trust high-cosine reads, corroborate specifics. Output is
**English-only by construction** (constrained decoding).

## Verified metrics (held-out, reproducible)

Metric = the paper's **FVE** (0 = noise floor) + cosine + retrieval@1.

- **AV generative round trip** (activation → text → activation): best-of-24
  cosine **0.28**, **retrieval@1 0.75 (~60× chance)**; 100% coherent English.
- **AR given-text reconstruction**: cosine **0.90**, retrieval@1 **0.99**, **FVE
  0.81**.
- **Out-of-distribution** (hand-written far-domain prompts): AV round trip
  cosine **0.70**, retrieval@1 **0.46**; **100% of readings are English.**

# What does this mean?

NLAs turn hidden states into readable text, showing models **think in themselves
before responding,** much like humans do with conscious and subconscious
thoughts.

**No,** this doesn't prove that AI models are or are not conscious. Science as a
whole has yet to put a single definition down for qualia, or consciousness as a
whole, meaning we cannot determine if GRaPE is conscious or not.

***

# Notes

- Anthropic's NLA research paper: https://www.anthropic.com/research/natural-language-autoencoders
- This is not a complete model; it is an NLA for GRaPE 2.1.
- Updates and announcements are posted on [Skinnertopia](https://www.skinnertopia.com/) and this Hugging Face repository.

***

_GRaPE 2.1 Flash is developed under the [SLAI (Skinnertopia Lab for Artificial Intelligence)](https://www.skinnertopia.com/) brand and released under the Apache 2.0 license._