File size: 9,159 Bytes
393bb89
9f9fbec
393bb89
1bb3265
 
b3d11b8
4a79e5b
9f9fbec
 
6d68f94
 
4a79e5b
6d68f94
9f9fbec
 
b3d11b8
 
6d68f94
393bb89
1bb3265
6d68f94
 
 
 
 
 
393bb89
6d68f94
 
393bb89
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b3d11b8
393bb89
 
 
 
b3d11b8
393bb89
 
b3d11b8
1bb3265
b3d11b8
1bb3265
b3d11b8
4a79e5b
 
 
 
 
 
b3d11b8
 
 
1bb3265
6d68f94
 
1bb3265
 
 
4a79e5b
 
6d68f94
 
393bb89
b3d11b8
393bb89
b3d11b8
393bb89
b3d11b8
393bb89
b3d11b8
393bb89
 
b3d11b8
393bb89
b3d11b8
1bb3265
393bb89
1bb3265
b3d11b8
393bb89
b3d11b8
393bb89
b3d11b8
393bb89
 
 
b3d11b8
 
393bb89
 
 
 
 
9f9fbec
 
393bb89
 
 
b3d11b8
6d68f94
393bb89
 
 
 
 
 
 
 
 
 
 
 
6d68f94
b3d11b8
393bb89
b3d11b8
393bb89
 
 
 
b3d11b8
393bb89
b3d11b8
393bb89
b3d11b8
393bb89
b3d11b8
 
393bb89
b3d11b8
 
393bb89
 
 
 
 
 
 
 
 
b3d11b8
393bb89
 
 
 
 
 
 
 
b3d11b8
393bb89
b3d11b8
393bb89
b3d11b8
393bb89
6d68f94
393bb89
b3d11b8
6d68f94
393bb89
 
 
 
 
 
 
 
 
6d68f94
 
393bb89
 
 
6d68f94
393bb89
 
 
 
6d68f94
393bb89
6d68f94
393bb89
6d68f94
393bb89
6d68f94
1bb3265
b3d11b8
 
1bb3265
 
393bb89
 
 
 
 
 
 
1bb3265
 
 
 
 
 
b3d11b8
 
393bb89
6d68f94
393bb89
b3d11b8
393bb89
 
 
b3d11b8
393bb89
b3d11b8
393bb89
b3d11b8
393bb89
 
 
b3d11b8
1bb3265
393bb89
 
 
1bb3265
b3d11b8
393bb89
 
 
 
 
b3d11b8
393bb89
b3d11b8
393bb89
6d68f94
393bb89
 
 
 
 
 
 
 
 
 
 
 
 
 
b3d11b8
393bb89
4a79e5b
393bb89
4a79e5b
1bb3265
 
 
 
 
 
 
 
4a79e5b
1bb3265
 
393bb89
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
# FeatureLens v0.7 validation

This guide tests the **v0.7 changes** plus a small number of high-value regressions. Use the exact UI labels below. The comprehensive edge/adversarial suite remains deferred until final hardening.

## A. Local release gates

From the repository root:

```bash
python3 -m pytest -q
python3 -m compileall -q app.py featurelens experiments scripts
python3 scripts/ui_smoke.py
python3 scripts/release_check.py
```

Expected:

```text
45 passed
FeatureLens UI launch smoke: PASS
FeatureLens release check: PASS
  discovery prompts: 224
  causal tasks: 28
  layers: [4, 14, 26]
  feature-set sizes: [1, 3, 5]
  random controls: 8
  release: v0.7.0
```

## B. UI cleanup, table typography, focus mode, and export names

1. Open **Workbench**.
2. Confirm these are normal always-visible sections, **not collapsible dropdowns/accordions**:
   - **III. Single-feature scale dose-response**
   - **IV. Contrastive continuation preference**
3. Confirm there is only one experiment heading before each section. A plot title such as `Scale dose-response` is fine; there should not be a second accordion heading with the same text.
4. After any result table appears, confirm its table title (for example **Strongest active SAE features**) is visually larger/bolder than ordinary body text, and column headings are clearly readable.
5. Produce a plot and click its top-right **fullscreen/focus** icon.
   - The focused plot should open near the top-center of the viewport.
   - It should be bounded rather than stretching across the entire monitor.
   - Axis labels should remain readable without needing to pan horizontally across a huge plot.
6. Exit focus mode. You should return to approximately the same page location you were at before opening it.
7. Repeat the focus test on one wide result table.
8. Click the plot **export** icon. Verify the downloaded filename is descriptive, for example:

```text
featurelens_activation-profile.png
featurelens_scale-dose-response.png
```

It should not be the generic `chart.png`.
9. Click **Copy table with headers** and confirm the temporary `✓ Copied with headers` feedback still appears.

## C. Establish the shared Workbench context

In **Workbench → I. Inspect a prompt location** set:

```text
Prompt: The derivative of x squared is
Residual layer: 14
Prompt token index: -1
Displayed active features: 12
```

Click **Inspect sparse features**.

Expected regression values are approximately:

```text
selected token: 5 (' is')
feature 22632 activation: 32.90625
reconstruction cosine: 0.9015
NMSE: 0.1874
Top-5 activation mass: 26.7%
```

The persistent **Current Workbench context** banner must state the inspected prompt, layer 14, and token 5. It should remain visible when switching tabs.

## D. German-language control regression

Open **Feature evidence → A. Concept-guided candidate discovery** and inspect **Target concept**.

Pass conditions:

- `german_language` is available;
- `french_language` is not available.

The controlled examples in **Workbench** should also include a German example such as:

```text
Ich möchte einen Tisch für zwei reservieren.
```

No GPU run is required for this check.

## E. Balanced concept-guided candidate discovery

Open **Feature evidence → A. Concept-guided candidate discovery**.

Set:

```text
Target concept: mathematics
Residual layer: 14
Prompts per concept: 4
Candidate features: 12
Candidate ranking: Balanced selectivity
```

Keep the Workbench context from Test C and click **Discover concept-associated candidates**.

Verify that **Candidate feature evidence** includes these fields:

```text
Rank
Feature id
Candidate score
Target mean max
Other mean max
Mean difference
Selectivity
Target activation rate
Other activation rate
Current prompt max
Current token activation
Active at current token
```

Interpretation check:

- `Candidate score` is an exploratory balanced score, not a semantic label.
- Globally huge but weakly selective features should no longer automatically dominate merely because their raw coefficients are enormous.
- `Current prompt max` and `Current token activation` show whether a candidate can immediately be tested in the current Workbench context.
- **Selected candidate feature id** should default to the highest-ranked displayed candidate active at the current Workbench token when such a candidate exists; otherwise it may default to the top-ranked candidate.

Record the top 12 rows. These are the most scientifically useful new v0.7 results.

### E2. Raw-ranking comparison

Without changing the other fields, set:

```text
Candidate ranking: Raw mean difference
```

and rerun **Discover concept-associated candidates**.

This should reproduce the scale-sensitive style of ranking used previously. The purpose of this test is to verify that v0.7 exposes the methodological choice explicitly instead of silently changing it.

You do **not** need to investigate all 12 raw-ranking features causally.

## F. Candidate row selection and explicit reuse

After Test E:

1. Click a row in **Candidate feature evidence**. The row's feature id should appear in **Selected candidate feature id**.
2. Click **Use selected candidate across feature tests**.
3. A visible confirmation should say that the feature was loaded.
4. Verify the same id appears in:
   - **Workbench → Single feature id**;
   - **Workbench → Dose-response feature id**;
   - **Workbench → Contrastive feature id**;
   - **Feature evidence → Feature id**.

This is a state-routing operation and should not consume another GPU call.

## G. Completion-cue regression for feature 22632

Open **Feature evidence → C. Completion-cue sensitivity**.

Set:

```text
Feature id: 22632
Residual layer: 14
Prompt stem: The derivative of x squared
Completion cues:
is
=
:
equals
therefore
```

Click **Run completion-cue scan**.

The `is` condition reconstructs the original prompt and should remain approximately:

```text
activation: 32.90625
Active in TopK: True
```

The previous reference result had zero activation for `=`, `:`, `equals`, and `therefore`. Treat that as a regression reference, not as a hard scientific requirement if tiny numerical differences appear.

## H. Cue × context specificity

Open **Feature evidence → D. Cue × context specificity**.

Use:

```text
Feature id: 22632
Residual layer: 14

Prompt stems:
The derivative of x squared
The capital of Germany
The weather today
My name

Completion cues:
is
=
:
equals
therefore
```

Click **Run cue × context scan**.

Expected output:

- 20 rows = 4 stems × 5 cues;
- columns **Prompt stem**, **Cue**, **Full prompt**, **Final token**, **Activation**, **Active in TopK**;
- a **Cue response across contexts** plot.

Scientific question: does feature `22632` fire specifically on the literal `is`, on `is` across many contexts, on several completion-boundary cues, or only in the original mathematics context?

Send the active rows and the per-cue context summary. This is the other key new v0.7 result.

## I. Prompt-wide selected-feature contrast with German control

Still in **Feature evidence**, set:

```text
Feature id: 22632
Residual layer: 14
Prompts per concept: 4
```

Click **Run controlled concept contrast**.

The result should contain `german_language` rather than `french_language` while preserving seven balanced concepts and 28 sampled prompts total.

Do not expect the exact old French-based concept means to remain identical because one control group has intentionally changed.

## J. High-value existing regressions

Only confirm these run; do not resend full tables unless something changes unexpectedly.

1. **Workbench → III. Single-feature scale dose-response** with feature `22632`, target `2x`:
   - multiplier `1.0` must remain the exact same-batch numerical null with all causal deltas equal to zero.
2. **Paraphrase robustness** default pair:
   - selected-token cosine should remain around `0.301`;
   - prompt-wide cosine should remain around `0.981`.
3. **Paraphrase robustness** identical-prompt control:
   - all four robustness metrics should remain `1.000`.
4. **Layer trajectory**:
   - layer 14 should remain the weakest reconstruction of layers 4/14/26 on the default prompt.
5. **Feature sets → Set-size sensitivity → Run 1/3/5-feature ablation sweep**:
   - this is the top-k causal sweep;
   - it is **not** the Workbench `Strongest active SAE features` table;
   - confirm k = 1, 3, 5 rows appear.
6. After dynamic results expand, scroll to the actual bottom without switching tabs. The v0.5 reflow fix should remain intact.

## K. Deferred final hardening

Do not spend ZeroGPU quota on the complete adversarial suite yet. Preserve these for the final hardening release:

- empty prompts;
- out-of-range token indices;
- feature ids `0`, `32767`, invalid `32768`;
- inactive feature ablation vs injection;
- whitespace-sensitive targets such as `2x` vs ` 2x`;
- missing required target continuations;
- unrelated paraphrases;
- negation and mixed-language prompts;
- long-prompt truncation;
- rapid repeated clicks / queue behavior;
- narrow/mobile-width layout;
- final copy/focus/export sweep.