File size: 12,495 Bytes
0e55eb5 1acbf39 0e55eb5 1acbf39 0e55eb5 1acbf39 0e55eb5 1acbf39 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 | ---
title: Transformers Architecture Inspector
emoji: π
colorFrom: gray
colorTo: yellow
sdk: static
app_file: index.html
pinned: false
short_description: Inspect the semantic architecture IR of Transformers models
---
# Transformers Architecture Viewer
An interactive, web-based viewer for the **Transformers Architecture IR** β a
compact, semantic, config-parametric representation of a model architecture.
The viewer consumes the IR artifacts directly (it never inspects Transformers
Python code, never loads checkpoint weights, and never instantiates a model).
The graph layout is derived entirely from the IR: semantic components become
nodes, symbolic repeats become collapsible blocks, and coarse dataflow becomes
typed edges.
> This is a **UI prototype** for a standalone Hugging Face Space, not the final
> production Hub integration.
## What it shows
- **Semantic components** as nodes, colour-coded by kind (attention, feed-forward,
normalization, embedding, position, β¦).
- **Symbolic repeats** as collapsible blocks, labelled e.g. `Decoder Layer Γ 32`
(resolved against a config) or, in **config-agnostic** mode,
`Decoder Layer Γ num_hidden_layers` β expanding a repeat then draws a stacked
"deck" with a `β― repeated β―` hint to convey the fully-deployed model without
unrolling all N instances.
- **Hub checkpoints**: a dropdown lists the most-downloaded public models on the
Hugging Face Hub that use this architecture; picking one loads its real
`config.json` and re-resolves every repeat count live.
- **Hierarchy** as containment boxes and **coarse dataflow** as arrows, rendered
separately.
- **Edge kinds** as visually distinct arrows: `data`, `residual`, `mask`,
`position`, `cross_attention` β each toggleable.
- **Provenance**, per-component **attributes** (semantic facts), model-level
**architecture facts**, and config-derived values for the selected component.
- **On-node captions**: each node shows a shape/attribute line derived from the
IR β e.g. `[B, S, 4096]` (observed dataflow, config-resolved) or
`head dim 128 Β· n heads 32` (attributes). Toggle with *Show shapes & attributes*.
- **Colour-by-role** node fills (attention / feed-forward / normalization / β¦).
- **Manual repositioning**: drag any node/group to declutter edges.
### Comparison mode
The **β Compare** button opens a two-architecture diff. It picks a comparison
regime from the data and renders accordingly:
- **Step 0 β regime**: if one model `extends` the other (or `modularity`
names a parent) β *same lineage* (align by stable id). Otherwise it measures
stable-id overlap: high β *shared ids* (align by id); low β *cross-lineage*
(align by semantic **kind lanes**). A `modular_graph.json` forest file would
let this also suggest lineage comparisons β the code reads lineage from each
artifact's `extends`/`modularity` today.
- **Step 1 β headline**: the `architecture` facts side-by-side
(`decoder Β· MHA Β· rope` vs `enc_dec Β· MHA Β· relative`), diffs highlighted.
- **Step 3 β structure**: matched nodes' `attributes` diffed (e.g. GQA shows up
as `n_kv_heads 32 β 8`); added / removed nodes flagged. Cross-lineage pairs
fall back to kind lanes (counts + representative attribute deltas per kind).
- **Step 4 β scale**: hidden size, resolved depth (`repeats` Γ `config`),
intermediate size, vocab. Pick a Hub checkpoint per side to make it concrete.
- **Step 5 β topology**: edge-kind presence β surfaces `cross_attention`
(enc-dec) and `cache_*` (decoder KV cache) that attribute diffs miss.
**Side-by-side graphs**: both architectures render in synced panels (pan/zoom
move together). Components present in **both** models β matched by stable id β
get a green "shared" outline in each panel, so common structure pops out while
the differences stand alone. Below the graphs sits the full textual diff report.
The viewer owns all alignment + delta math; there is no shipped pairwise-diff.
Same-lineage overlay rendering (ghosted base + highlighted patches on a single
graph) is the next step; today the side-by-side view already highlights shared
components. Current samples (llama/bert/t5) are standalone and cross-family, so
they share few components β the shared outline lights up for lineage pairs
(e.g. gemma vs llama) once those artifacts exist.
### How close can it get to a hand-drawn architecture poster?
The *visual language* (nested dashed containers, class-name headers, `Γ N`
badges, colour-by-role, shape/attribute captions, config-resolved dims) is fully
IR-driven and rendered generically. Two things in a richly hand-drawn diagram
are **not** reachable from the current IR without either deepening the generator
or hardcoding one architecture:
- **Sub-module detail** β `q/k/v/o_proj`, `gate/up/down_proj`, the SDPA box:
`LlamaAttention` / `LlamaMLP` are leaves (`children: []`) in the IR. If the
generator descends into them (and annotates dims/attributes), the viewer shows
them automatically β no viewer change needed.
- **Task-head / runtime extras** β LM head, softmax, logits, tokenizer chips,
the causal-mask heatmap: the artifact is the base model (`LlamaModel`), and
tokens/masks are runtime, not architecture. These would need new IR content
(e.g. a `β¦ForCausalLM` artifact) or a separate data source.
### IR spec
The viewer conforms to the current `architecture-template-v0` artifacts (as
emitted by the generator into `ir/`). It does not carry any legacy shims β
regenerate the artifacts and the viewer follows.
- **Config block**: reads `config.referenced_fields` (authoritative parametric
surface) + `config.salient_fields` (curated scalars), merged (referenced wins)
to drive the resolver and the editable config. The full `config.fields` blob
is intentionally no longer produced or read.
- **Edge kinds** are discovered from the artifact, so `cache_read` / `cache_write`
(and any future kind) get a colour, an arrowhead and a legend toggle. Kinds
not in the known set fall back to a deterministic palette colour.
- **Pseudo endpoints**: `input:*` (e.g. attention mask) and `state:*` (e.g. the
KV cache) render as distinct source/sink pills so cache-read/write arrows have
visible endpoints.
- **Semantic facts**: model-level `architecture` (family / view / positional /
attention variant / MoE) shows as a sidebar summary and on the model node;
per-component `attributes` and the `moe` kind appear in the inspector.
- **Observed dataflow**: the `dataflow` block's per-stage symbolized shapes
(e.g. `[B, S, config.hidden_size]`) are shown in the inspector β whole-model
input/output on the model node, and in/out shapes on each captured component.
- **Modular inheritance**: when `modularity.is_modular` is set, the sidebar
shows the parent model, patch count and `diff_size`.
## Tech choices
Zero-build, dependency-free **static site** (vanilla JS + SVG):
- No npm install, no bundler, no CDN β so it is CSP-safe and trivial to host on a
`sdk: static` Space.
- Runs by serving the folder with any static file server.
- The layout engine (`js/layout.js`) is a small nested/layered graph layout; the
resolver (`js/ir.js`) is a tiny client-side evaluator for repeat count
expressions such as `config.num_hidden_layers`.
## Run locally
Any static server works (a server is needed because the app uses `fetch` +
ES modules, which browsers block on `file://`):
```bash
# from the repository root
python3 -m http.server 8000
# then open http://localhost:8000
```
or
```bash
npx serve . # http://localhost:3000
```
## Deploy as a Hugging Face Space
One command pushes the app + everything under `ir/` to a static Space
(creates it on first run, then re-uploads on each call):
```bash
./deploy.sh # β lysandre/transformers-architecture-inspector
./deploy.sh <owner>/<space-name> # β your own Space
```
Requires the `hf` CLI (`pip install -U huggingface_hub`) and a logged-in session
(`hf auth login`). The script runs `hf repo create --type space --space-sdk static`
then `hf upload . .` (excluding `.git`, `.idea`, `deploy.sh`). The Space serves
`index.html` at the root and the app fetches `ir/manifest.json` + its artifacts.
## How it loads IR artifacts
The app reads the generator's output directly β **`ir/` is the single source of
truth, there is no copy step**. Point your generator's export at `ir/` and a
reload picks up the new files.
- On load it fetches **`ir/manifest.json`** and populates the **architecture
selector** from `manifest.architectures[]` (currently `llama`, `bert`, `t5`).
- Manifest artifact paths (e.g. `artifacts/llama.json`) are resolved relative to
the manifest, i.e. `ir/artifacts/llama.json`.
- **Hierarchy reconstruction**: most artifacts link submodules via `children`,
but some (multimodal models like gemma3/internvl/qwen2_5_vl/siglip2) leave
`model.children` empty and encode the tree only in dotted `path_pattern`s. The
viewer detects orphaned components/repeats and reattaches them by path,
synthesizing intermediate containers (`vision_tower`, `language_model`, β¦) so
every model is spreadable. Well-linked models are untouched.
- You can also **upload** an IR JSON file or **paste** artifact JSON directly.
- You can **paste / edit a config** (the merged referenced + salient fields) and
click *Apply & resolve* to re-resolve all symbolic repeat counts live.
- Or pick a **Hub checkpoint** from the dropdown to fetch its `config.json`
(both the Hub models API and `resolve/main/config.json` are called directly
from the browser β no backend; they return permissive CORS headers). Gated or
private checkpoints report a clear message instead of loading.
### Where to export from the generator
Write straight to **`ir/`** (its current output dir):
```
ir/manifest.json # architecture index
ir/artifacts/<model>.json # one artifact per architecture
```
The `IR_BASE` constant at the top of `js/app.js` controls this location; change
it there if you want the app to read from somewhere else.
## Project layout
```
index.html # app shell: header, sidebar, canvas, inspector
style.css # Hugging Face-flavoured styling
js/ir.js # IR model + client-side repeat-count resolver
js/layout.js # nested layered graph layout derived from the IR
js/graph.js # SVG renderer + pan/zoom/select/collapse
js/app.js # wiring: loading, sidebar controls, inspector (IR_BASE here)
ir/manifest.json # architecture index β the generator writes here
ir/artifacts/*.json # IR artifacts β the app reads these directly
```
## Controls
- **Scroll** to zoom, **drag the canvas** to pan, **Fit** to recenter.
- **Drag a node** to reposition it (dragging a group box moves its whole
subtree); edges follow live. **Reset layout** clears manual positions.
- **β** toggles light / dark theme (defaults to the HF dark palette).
- **Click** a node to inspect it; click an edge entry in the inspector to jump.
- **β / β** on a repeat block expands / collapses it (repeats are **expanded by
default**). *Collapse / Expand repeats* toggle all at once.
- Sidebar: architecture selector, Hub-checkpoint dropdown, config editor,
edge-kind toggles, config-agnostic and provenance options.
## Known limitations (prototype)
- Layout is a lightweight custom layered algorithm; very large/expanded graphs
are readable but not crossing-optimal (no ELK/Sugiyama ordering pass).
- Expanded repeats render one representative body instance (plus a stacked-deck
hint in config-agnostic mode), not N unrolled copies β matching the IR.
- Edge routing is straight bezier (no orthogonal routing / obstacle avoidance).
- Non-repeat containers are always expanded; only symbolic repeats collapse.
- The Hub dropdown filters by the `model_type` tag, top 30 by downloads β not an
exhaustive search; gated checkpoints can't expose their config to the browser.
- No search / filter across components yet.
## Next steps
- Optional ELK.js layout backend for large expanded graphs.
- Component search + "focus subtree" navigation.
- Deep-link state (selected node, expand set, config) in the URL.
- Hub integration: load the IR artifact itself by `model_id` (today only the
`config.json` is fetched from the Hub; the IR still comes from `ir/`).
- Diff view between two configs (e.g. base vs. resolved layer counts).
|