FeatureLens / CHANGELOG.md
ArchitSharma's picture
Publish FeatureLens study
ffa621b
|
Raw
History Blame Contribute Delete
23.9 kB

A newer version of the Gradio SDK is available: 6.25.0

Upgrade

Changelog

v1.0.0

  • Published the completed offline study artifacts and public Study dashboard.
  • Finalized causal inference around two position policies: final prompt token and maximum selected-feature activation within the prompt.
  • Reported causal-task-level paired inference, separating intervention coverage from conditional effect strength.
  • Measured selected SAE features at 0.962 mean held-out AUROC; max-active causal edits reached 82.1% task coverage and 2.33× matched-random target effect on average.
  • Kept the weaker final-token baseline as a position-sensitivity control rather than replacing it.
  • Added the committed reproducibility bundle, including both Colab notebooks under notebooks/, fixed split metadata, measured CSV/JSON outputs, report, and figures.
  • Simplified release configuration/checks by removing historical per-version feature bookkeeping from research_config.json.
  • Public interface remains version-neutral; release metadata is internal to the repository.

v0.16.0

  • Added two explicit offline causal-position policies: final_token and max_feature_activation. Max-active positions are selected only from SAE activation within the prompt, never from behavioral outcomes.
  • Preserved the v0.15 final-token causal result as a positional baseline and added a causal-addendum runner that computes only the new max-active policy.
  • Changed primary paired causal inference to use the causal task as the statistical unit, averaging ablation and 2× amplification within each task before bootstrap/sign-flip inference.
  • Added exact sign-flip enumeration for small effective paired samples, with deterministic Monte-Carlo fallback for larger samples.
  • Separated causal coverage from conditional-on-active effect strength and added final-token vs max-active position-sensitivity summaries.
  • Added causal_position_summary.csv, a position-sensitivity report figure, max-active association-vs-causality synthesis, and updated Study-tab diagnostics.
  • Added experiments/run_causal_addendum.py and a Drive-backed Colab addendum notebook so a completed v0.15 study can be upgraded without recollecting discovery activations or rerunning feature-set inference.
  • Updated artifact validation and methodology documentation for the finalized v0.16 study schema.

v0.15.0

  • Reworked the public Gradio surface around a documented research-instrument design system rather than SaaS/dashboard defaults.
  • Added DESIGN.md with typography, color, spacing, surface, button, table, plot, and anti-pattern rules so future UI edits have explicit constraints.
  • Replaced the three-column onboarding/card pattern with compact editorial guidance; flattened repeated context/callout treatment; removed visible release marketing; shortened result copy so measured values lead and methodology lives in the Method tab/docs.
  • Introduced a two-typeface hierarchy (serif display headings, neutral sans-serif controls/data), tighter semantic spacing, compact primary actions, quiet utility buttons, and stronger explicit table headings.
  • Normalized the dynamic cross-target plots to a restrained three-series palette instead of Vega's saturated categorical defaults.
  • Rewrote the public README around the research question, live tool, offline study, and reproducible workflow instead of a long release-history narrative.
  • Added a ready-to-run Google Colab notebook plus docs/COLAB.md for Drive-backed artifact persistence and resumable study execution.
  • Added --activation-batch-size / --activation-max-length to the full runner and task-level checkpoint/resume support inside causal and feature-set stages.
  • Added automated design-contract regression tests.

v0.14.0

  • Transitioned the project from live-feature expansion toward the full offline empirical study.
  • Changed offline SAE concept evidence from final-token-only activations to prompt-wide max-pooled activations across non-padding tokens, while saving separate final-token sparse activation matrices for local diagnostics.
  • Added experiments/analyze_stability.py with 128 deterministic balanced activation resamples and per-feature shortlist support/rank summaries.
  • Added experiments/analyze_study.py to join held-out AUROC/F1, paraphrase robustness, candidate stability, feature activity, and random-normalized target/JS causal specificity by controlled concept.
  • Added descriptive cross-concept association-vs-causality correlations and new association/candidate-stability report figures.
  • Expanded the public Offline study tab into an artifact-backed results dashboard that remains explicitly empty until real study artifacts are committed.
  • Added python experiments/run_all.py --resume for interrupted/preemptible GPU sessions and python experiments/run_analysis_only.py for CPU-only re-analysis once inference artifacts exist.
  • Added scripts/validate_artifacts.py to verify prompt-wide activation provenance and public study artifact schemas before commit.
  • Added dedicated offline-study methodology/validation documentation and automated tests for prompt-wide pooling, stability scoring, task-paired specificity, study UI states, and correlation guardrails.

v0.13.0

  • Added deterministic 32-resample candidate-support diagnostics from the same concept-discovery activation batch; each displayed feature now reports shortlist support and median resample rank without another model forward.
  • Extended cross-target profiling with normalized effect entropy, effect concentration, signed bias, and a descriptive profile pattern so concentrated target dependence is separated from broad same-sign behavior.
  • Added pairwise target-preference shifts derived from the same cross-target scores: Δ(A−B) = Δmean(A) − Δmean(B), with a table and plot and no additional inference.
  • Kept HF acceptance quota-aware: only concept discovery and cross-target profiling are touched GPU paths.

v0.12.0

  • Added split-half discovery stability using the already-computed concept activation batch, so shortlist sensitivity is visible without another GPU forward.
  • Added Controlled evidence patterns, a zero-GPU synthesis that distinguishes broad controlled influence, target-weighted effects, distribution-shift-dominant effects, and weak/mixed specificity while keeping eight-control tails explicitly coarse.
  • Changed Association vs controlled causality so missing discovery state after a Space rebuild produces an explicit explanation instead of a blank panel.
  • Added Cross-target causal profile for up to three candidates and five exact continuations, screening whether native ablation effects concentrate on one target or generalize across alternatives.
  • Added automatic cross-target shortlist handoff from the target-specificity and JS-specificity leaders.
  • Kept the validated in-place focus behavior unchanged.
  • Kept HF validation quota-aware: only the touched discovery path and the new cross-target path need live GPU acceptance.

v0.11.0

  • Added Controlled candidate specificity, a one-batch follow-up that compares up to three candidate SAE ablations against each candidate's own 8-direction norm-matched random ensemble.
  • Added a strategic controlled shortlist that preserves the discovery leader, target-effect leader, and distribution-shift leader when they differ, then fills remaining slots by triage target rank.
  • Added target-specificity and JS-specificity ratios plus coarse empirical random-control tails for every controlled candidate.
  • Added Association vs controlled causality, joining concept-discovery evidence to random-normalized causal specificity instead of relying only on raw triage magnitude.
  • Kept target-specific and whole-distribution causal influence separate rather than collapsing them into one score.
  • Kept the validated in-place focus/zoom implementation unchanged.
  • Reduced HF acceptance to one new GPU call; unchanged discovery/triage and other regression paths remain covered by automated tests.

v0.10.0

  • Added a zero-extra-GPU Discovery–causality alignment panel after candidate triage.
  • Joined discovery rank/score with target-effect rank and next-token-distribution-shift rank for the same screened candidates.
  • Added descriptive Spearman ρ summaries for candidate evidence vs absolute target effect and vs next-token JS.
  • Added an association-evidence-vs-target-effect scatter plot and rank-shift table.
  • Kept the v0.9 in-place focus implementation unchanged after HF validation.
  • HF validation remains GPU-budget-aware: only the discovery and candidate-triage paths need to be exercised.

v0.9.0

In-place focus and layout polish

  • Replaced the HF-iframe-hostile overlay/fullscreen experiment with in-place focus for plots and tables. The original component expands exactly where it is located; plots are scaled from their existing rendering so aspect ratio is preserved and no cloned toolbar icon can be mistaken for the chart.
  • Focus is capped by both screen width and height, and the same fullscreen toolbar icon toggles the component back without moving the page to the top.
  • Reworked explicit result-table headings into compact HTML headings that occupy the Dataframe toolbar whitespace instead of leaving a large empty band above the first row.
  • Removed the unnecessary “Standalone experiment” dose-response explanation while keeping independent feature and target fields.

Candidate-to-causality workflow

  • Added Batched causal candidate triage. Up to eight concept-discovery candidates are independently ablated in one batched scoring run at the current Workbench location.
  • The screen reports native activation, perturbation norm, target mean/sequence log-probability deltas, and next-token JS, and ranks candidates by absolute target effect.
  • The triage deliberately omits random controls; its purpose is to identify which candidate is worth promoting to the existing single-feature 8-direction random-control test.
  • Concept discovery now directly populates the candidate-screen multiselect, defaulting to up to five returned candidates.

GPU-budget-aware validation

  • HF acceptance no longer reruns identity paraphrase, layer trajectory, or 1/3/5 set-size sweeps when those code paths are unchanged. Automated tests cover them; scarce ZeroGPU minutes are reserved for new/touched inference paths.

v0.8.0

UI readability and focus

  • Replaced plot-native fullscreen behavior with a bounded FeatureLens focus overlay. The plot is copied into a centered reading surface (max ~1120 px) instead of stretching across an ultrawide display; closing the overlay restores the original page position.
  • Kept descriptive plot export filenames and the Gradio Dataframe fullscreen control.
  • Replaced fragile native Dataframe labels with explicit result-table headings above every major table so table titles follow the same typography hierarchy as the rest of the application.
  • Added a muted cue palette to the cue × context plot instead of relying on Gradio/Vega default saturated series colors.

Independent experiment inputs

  • The scale dose-response panel now owns its Dose-response feature id and Dose-response target continuation. It no longer depends on running the single-feature causal test or filling that section's optional target field first.
  • Clarified in-panel provenance: dose response reads the current prompt/layer/token fields from Workbench Section I but is otherwise a standalone experiment.

Candidate discovery

  • Added Causal-ready at current token ranking. It requires positive concept contrast and activation at the selected Workbench token, then ranks those compatible candidates using balanced selectivity plus a log-scaled current-token activation term.
  • Discovery summaries now report how many displayed candidates are actually active at the selected Workbench token. This makes the distinction between a prompt-wide concept candidate and an immediately ablatable feature explicit.
  • Balanced selectivity and raw mean-difference modes remain available for methodological comparison.

Cue specificity

  • Cue × context summaries now derive the dominant cue, its context coverage, and off-dominant activity. A feature that fires for one cue in every tested context while all other cues stay inactive is reported as a cue-dominant tested pattern, not merely with generic interpretation text.

Validation

  • Expanded the automated suite to cover causal-ready candidate discovery, independent dose-response target state, cue-dominance diagnostics, plot-focus JavaScript markers, and explicit result-heading behavior.
  • Retained compile, Ruff, actual Gradio launch(), release-check, and deferred final adversarial-suite gates.

v0.7.0

Research-instrument UI cleanup

  • Removed collapsible wrappers from the core scale dose-response and contrastive preference experiments so section headings are no longer duplicated by accordion titles.
  • Simplified the page header and removed visible footer/redundancy that did not help a reviewer use the tool.
  • Strengthened result-table title and column-header typography.
  • Replaced full-viewport plot stretching with a bounded top-centered focus view; exiting focus restores the prior page position.
  • Kept native plot export but rename downloads to descriptive featurelens_<plot-name>.png filenames instead of a generic chart name.

Candidate discovery and causal readiness

  • Reworked live concept-guided discovery around Balanced selectivity: selectivity × target activation rate × log1p(target mean). This prevents very large but non-selective SAE coefficients from dominating the exploratory shortlist.
  • Retained Raw mean difference as an explicit comparison mode rather than silently changing the old ranking.
  • Candidate discovery now evaluates the current Workbench prompt in the same GPU batch and reports current-prompt maximum activation plus selected-token activation.
  • The candidate selector defaults to the highest-ranked displayed candidate active at the current Workbench token when one exists.
  • Candidate-table row selection is wired directly to the candidate selector, and the reuse action now confirms exactly which downstream feature selectors were updated.

Lexical / structural specificity

  • Added a cue × context specificity matrix: cross several prompt stems with the same completion cues in one batched forward and measure the selected feature at every resulting final token.
  • This extends the single-stem completion-cue test so a response to is can be separated from a broader completion-boundary or context-dependent response.

Controlled data and validation

  • Replaced the French-language control concept one-for-one with a German-language concept while preserving 224 balanced discovery prompts and 28 causal tasks.
  • Added regression coverage for German data, balanced/raw candidate ranking, current-Workbench candidate compatibility, cue × context scanning, candidate row selection, descriptive export naming, and focus-position preservation.
  • Kept the actual Gradio launch() smoke test, compile gate, release checker, and deferred comprehensive adversarial suite.

v0.6.0

UX / navigation

  • Added a plain-language Start here tab with a three-step workflow and glossary for non-specialist reviewers.
  • Added a persistent Current Workbench context banner so inherited prompt/layer/token state is visible from every tab.
  • Added explicit editable feature selectors for scale dose-response and contrastive continuation preference instead of silently reusing the single-feature selector.
  • Clarified state provenance in Feature Sets and Feature Evidence; experiment text now states whether it inherits Workbench context or uses an independent prompt set.
  • Normalized heading hierarchy and increased table/header typography for readability.
  • Added native plot fullscreen and export PNG controls to every BarPlot/LinePlot.

Live research tools

  • Added concept-guided candidate feature discovery: rank features for a selected controlled concept using prompt-wide target-minus-other mean maximum activation, with selectivity and activation-rate diagnostics.
  • Added a completion-cue sensitivity scan that appends controlled suffixes to a prompt stem and measures the selected feature at the resulting final token.
  • Added a one-click action to reuse a discovered candidate across single-feature, dose-response, contrastive, and evidence feature selectors.
  • Candidate discovery and cue scans are explicitly exploratory; neither creates semantic labels or overwrites Offline concept hint.

Validation

  • Expanded toy-runtime coverage to 40 tests, including candidate-feature discovery and completion-cue scans.
  • Retained the actual Gradio launch() smoke gate, compile gate, release checker, batched-null regression tests, and final-release adversarial-test deferral.

v0.5.0

Causal specificity

  • Added a contrastive continuation preference test that scores two exact continuations under the same single-feature intervention and 8-direction norm-matched control ensemble.
  • Reports baseline/edited sequence log-odds A−B, causal log-odds shift, token-normalized preference shift, random-control magnitude statistics, and an exploratory empirical tail probability.
  • Keeps this distinct from absolute target probability so broad distributional disruption is not mistaken for selective behavioral control.

Feature evidence and geometry

  • Added a feature-token activation trace over every token in the current Workbench prompt.
  • Fixed the controlled concept scan to use prompt-wide max activation over non-padding tokens instead of only the final token.
  • All-zero concept batches now report inactive in every sampled prompt and do not invent a leading concept.
  • Added feature-set decoder geometry for 2–8 selected features: pairwise decoder cosine, mean/max absolute cosine, activation-weighted joint-ablation norm, independent-direction reference norm, and alignment/cancellation ratio.

Interface

  • Widened and explicitly centered the application canvas (up to 1600 px) and enabled fill_width=True to use desktop space more effectively.
  • Normalized serif typography, labels, controls, table font sizes, and action-button styling.
  • Added bounded Dataframe heights to reduce excessive dynamic page growth.
  • Added copy-button visual acknowledgement (✓ Copied with headers).
  • Added a browser-side resize/mutation observer to request layout reflow when dynamic output height changes inside an embedded Space.
  • Retained the safe Gradio theme configuration without string font tuples; serif typography is applied in CSS.

Validation

  • Expanded automated coverage from 29 to 38 tests, including contrastive log-odds, decoder geometry, copy/export helpers, and toy-runtime end-to-end checks.
  • Added scripts/ui_smoke.py so the actual Gradio launch() path is part of the release procedure instead of only constructing the component tree.
  • Reworked v0.5 acceptance tests around the new concept-scan semantics, feature-token trace, contrastive preference, geometry, copy feedback, and embedded-page reflow.

v0.4.0

Causal correctness

  • Added an explicit batched zero-edit reference to live causal batches so intervention effects are measured against the same execution context as edited rows.
  • Changed the scale dose-response reference to the batched row. The row is therefore an exact causal no-op by construction rather than a separately executed baseline comparison.
  • Added execution-context drift diagnostics so any remaining single-forward vs batched-forward numerical difference is reported as instrumentation drift, not causal signal.

Stronger negative controls

  • Replaced the single live random residual direction with an 8-direction norm-matched random ensemble.
  • Single-feature, joint feature-set, and 1/3/5 set-size experiments now report random signed mean, mean absolute effect, standard deviation, targeted/random magnitude ratio, and a small-sample empirical tail probability.
  • Updated offline causal and feature-set runners to use the same zero-edit reference and configurable random-control ensembles.
  • Updated report pairing so each targeted intervention is compared with the mean absolute effect of its complete random-control ensemble rather than an arbitrary first control.

Distributed causality

  • Added individual-vs-joint ablation decomposition for 2–5 features.
  • Reports the individual effects, additive expectation, observed joint effect, interaction excess, and normalized non-additivity.
  • The UI explicitly treats non-additivity as a diagnostic, not proof of a direct feature-feature circuit.

Representation robustness and evidence

  • Added prompt-wide paraphrase robustness using max activation per SAE feature across all prompt tokens, alongside the stricter selected-token comparison.
  • Added a live controlled concept contrast scan over the seven balanced discovery concepts for one selected feature.
  • Concept contrast results remain exploratory and never overwrite Offline concept hint or claim a semantic label from the live scan.

UI and export

  • Reworked the visual language toward a restrained, print-inspired interface with serif typography, flatter controls, thin rules, and muted chart colors.
  • Replaced intervention radio pills with conventional dropdown controls and realigned the Feature Sets form.
  • Added extra bottom spacing plus an explicit end-of-workbench footer to avoid an app-controlled abrupt cutoff in embedded Spaces.
  • Added Copy table with headers actions for every major output table; copied text is tab-separated and begins with the column names.
  • Added a dedicated Feature evidence tab for controlled feature-concept contrast tests.

Validation

  • Expanded the software suite to 29 tests.
  • Added report tests that verify random-control ensembles are aggregated correctly before paired causal statistics are computed.
  • Rewrote docs/VALIDATION.md around exact v0.4 UI labels, including explicit numerical-null, table-copy, adversarial, responsive-layout, and queue tests.

v0.3.0

Causal measurement

  • Replaced first-token-only target evaluation with exact full-continuation teacher-forced log-probability scoring.
  • Added total sequence and mean-per-token log-probability deltas plus per-target-token decomposition.
  • Retained next-token probability/JS diagnostics and greedy generation as complementary outputs.

Distributed feature causality

  • Added joint multi-feature ablation/scaling using summed reconstruction-preserving SAE decoder deltas.
  • Added live top-1/top-3/top-5 joint-ablation sweep with norm-matched random controls.
  • Added offline experiments/run_feature_sets.py and report integration.

Robustness

  • Added a live paraphrase-robustness explorer with TopK Jaccard, sparse cosine, overlap table, and activation comparison.

Efficiency

  • Batched all six single-feature dose-response edits into one model forward after the baseline.
  • Batched targeted/random 1/3/5 feature-set sweep conditions into one model forward after the baseline.

UI / deployment

  • Replaced the bright blue visual emphasis with muted teal/stone accents and explicit chart palettes.
  • Added visible Prompt tokens headings and aligned validation terminology with actual UI labels.
  • Explicitly labels the dose-response panel as a scale intervention: 0× = ablation, 1× = no edit.
  • Kept SSR disabled for the Hugging Face Space.

v0.2.0

  • Added live norm-matched random controls.
  • Added single-feature causal dose-response.
  • Added layer trajectory diagnostics.
  • Added bootstrap confidence intervals and paired sign-flip tests.
  • Hardened Gradio / ZeroGPU deployment.