File size: 12,495 Bytes
0e55eb5
1acbf39
 
 
 
0e55eb5
1acbf39
0e55eb5
1acbf39
0e55eb5
 
1acbf39
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
---
title: Transformers Architecture Inspector
emoji: πŸ”
colorFrom: gray
colorTo: yellow
sdk: static
app_file: index.html
pinned: false
short_description: Inspect the semantic architecture IR of Transformers models
---

# Transformers Architecture Viewer

An interactive, web-based viewer for the **Transformers Architecture IR** β€” a
compact, semantic, config-parametric representation of a model architecture.

The viewer consumes the IR artifacts directly (it never inspects Transformers
Python code, never loads checkpoint weights, and never instantiates a model).
The graph layout is derived entirely from the IR: semantic components become
nodes, symbolic repeats become collapsible blocks, and coarse dataflow becomes
typed edges.

> This is a **UI prototype** for a standalone Hugging Face Space, not the final
> production Hub integration.

## What it shows

- **Semantic components** as nodes, colour-coded by kind (attention, feed-forward,
  normalization, embedding, position, …).
- **Symbolic repeats** as collapsible blocks, labelled e.g. `Decoder Layer Γ— 32`
  (resolved against a config) or, in **config-agnostic** mode,
  `Decoder Layer Γ— num_hidden_layers` β€” expanding a repeat then draws a stacked
  "deck" with a `β‹― repeated β‹―` hint to convey the fully-deployed model without
  unrolling all N instances.
- **Hub checkpoints**: a dropdown lists the most-downloaded public models on the
  Hugging Face Hub that use this architecture; picking one loads its real
  `config.json` and re-resolves every repeat count live.
- **Hierarchy** as containment boxes and **coarse dataflow** as arrows, rendered
  separately.
- **Edge kinds** as visually distinct arrows: `data`, `residual`, `mask`,
  `position`, `cross_attention` β€” each toggleable.
- **Provenance**, per-component **attributes** (semantic facts), model-level
  **architecture facts**, and config-derived values for the selected component.
- **On-node captions**: each node shows a shape/attribute line derived from the
  IR β€” e.g. `[B, S, 4096]` (observed dataflow, config-resolved) or
  `head dim 128 Β· n heads 32` (attributes). Toggle with *Show shapes & attributes*.
- **Colour-by-role** node fills (attention / feed-forward / normalization / …).
- **Manual repositioning**: drag any node/group to declutter edges.

### Comparison mode

The **⇄ Compare** button opens a two-architecture diff. It picks a comparison
regime from the data and renders accordingly:

- **Step 0 β€” regime**: if one model `extends` the other (or `modularity`
  names a parent) β†’ *same lineage* (align by stable id). Otherwise it measures
  stable-id overlap: high β†’ *shared ids* (align by id); low β†’ *cross-lineage*
  (align by semantic **kind lanes**). A `modular_graph.json` forest file would
  let this also suggest lineage comparisons β€” the code reads lineage from each
  artifact's `extends`/`modularity` today.
- **Step 1 β€” headline**: the `architecture` facts side-by-side
  (`decoder Β· MHA Β· rope` vs `enc_dec Β· MHA Β· relative`), diffs highlighted.
- **Step 3 β€” structure**: matched nodes' `attributes` diffed (e.g. GQA shows up
  as `n_kv_heads 32 β†’ 8`); added / removed nodes flagged. Cross-lineage pairs
  fall back to kind lanes (counts + representative attribute deltas per kind).
- **Step 4 β€” scale**: hidden size, resolved depth (`repeats` Γ— `config`),
  intermediate size, vocab. Pick a Hub checkpoint per side to make it concrete.
- **Step 5 β€” topology**: edge-kind presence β€” surfaces `cross_attention`
  (enc-dec) and `cache_*` (decoder KV cache) that attribute diffs miss.

**Side-by-side graphs**: both architectures render in synced panels (pan/zoom
move together). Components present in **both** models β€” matched by stable id β€”
get a green "shared" outline in each panel, so common structure pops out while
the differences stand alone. Below the graphs sits the full textual diff report.

The viewer owns all alignment + delta math; there is no shipped pairwise-diff.
Same-lineage overlay rendering (ghosted base + highlighted patches on a single
graph) is the next step; today the side-by-side view already highlights shared
components. Current samples (llama/bert/t5) are standalone and cross-family, so
they share few components β€” the shared outline lights up for lineage pairs
(e.g. gemma vs llama) once those artifacts exist.

### How close can it get to a hand-drawn architecture poster?

The *visual language* (nested dashed containers, class-name headers, `Γ— N`
badges, colour-by-role, shape/attribute captions, config-resolved dims) is fully
IR-driven and rendered generically. Two things in a richly hand-drawn diagram
are **not** reachable from the current IR without either deepening the generator
or hardcoding one architecture:

- **Sub-module detail** β€” `q/k/v/o_proj`, `gate/up/down_proj`, the SDPA box:
  `LlamaAttention` / `LlamaMLP` are leaves (`children: []`) in the IR. If the
  generator descends into them (and annotates dims/attributes), the viewer shows
  them automatically β€” no viewer change needed.
- **Task-head / runtime extras** β€” LM head, softmax, logits, tokenizer chips,
  the causal-mask heatmap: the artifact is the base model (`LlamaModel`), and
  tokens/masks are runtime, not architecture. These would need new IR content
  (e.g. a `…ForCausalLM` artifact) or a separate data source.

### IR spec

The viewer conforms to the current `architecture-template-v0` artifacts (as
emitted by the generator into `ir/`). It does not carry any legacy shims β€”
regenerate the artifacts and the viewer follows.

- **Config block**: reads `config.referenced_fields` (authoritative parametric
  surface) + `config.salient_fields` (curated scalars), merged (referenced wins)
  to drive the resolver and the editable config. The full `config.fields` blob
  is intentionally no longer produced or read.
- **Edge kinds** are discovered from the artifact, so `cache_read` / `cache_write`
  (and any future kind) get a colour, an arrowhead and a legend toggle. Kinds
  not in the known set fall back to a deterministic palette colour.
- **Pseudo endpoints**: `input:*` (e.g. attention mask) and `state:*` (e.g. the
  KV cache) render as distinct source/sink pills so cache-read/write arrows have
  visible endpoints.
- **Semantic facts**: model-level `architecture` (family / view / positional /
  attention variant / MoE) shows as a sidebar summary and on the model node;
  per-component `attributes` and the `moe` kind appear in the inspector.
- **Observed dataflow**: the `dataflow` block's per-stage symbolized shapes
  (e.g. `[B, S, config.hidden_size]`) are shown in the inspector β€” whole-model
  input/output on the model node, and in/out shapes on each captured component.
- **Modular inheritance**: when `modularity.is_modular` is set, the sidebar
  shows the parent model, patch count and `diff_size`.

## Tech choices

Zero-build, dependency-free **static site** (vanilla JS + SVG):

- No npm install, no bundler, no CDN β€” so it is CSP-safe and trivial to host on a
  `sdk: static` Space.
- Runs by serving the folder with any static file server.
- The layout engine (`js/layout.js`) is a small nested/layered graph layout; the
  resolver (`js/ir.js`) is a tiny client-side evaluator for repeat count
  expressions such as `config.num_hidden_layers`.

## Run locally

Any static server works (a server is needed because the app uses `fetch` +
ES modules, which browsers block on `file://`):

```bash
# from the repository root
python3 -m http.server 8000
# then open http://localhost:8000
```

or

```bash
npx serve .        # http://localhost:3000
```

## Deploy as a Hugging Face Space

One command pushes the app + everything under `ir/` to a static Space
(creates it on first run, then re-uploads on each call):

```bash
./deploy.sh                                  # β†’ lysandre/transformers-architecture-inspector
./deploy.sh <owner>/<space-name>             # β†’ your own Space
```

Requires the `hf` CLI (`pip install -U huggingface_hub`) and a logged-in session
(`hf auth login`). The script runs `hf repo create --type space --space-sdk static`
then `hf upload . .` (excluding `.git`, `.idea`, `deploy.sh`). The Space serves
`index.html` at the root and the app fetches `ir/manifest.json` + its artifacts.

## How it loads IR artifacts

The app reads the generator's output directly β€” **`ir/` is the single source of
truth, there is no copy step**. Point your generator's export at `ir/` and a
reload picks up the new files.

- On load it fetches **`ir/manifest.json`** and populates the **architecture
  selector** from `manifest.architectures[]` (currently `llama`, `bert`, `t5`).
- Manifest artifact paths (e.g. `artifacts/llama.json`) are resolved relative to
  the manifest, i.e. `ir/artifacts/llama.json`.
- **Hierarchy reconstruction**: most artifacts link submodules via `children`,
  but some (multimodal models like gemma3/internvl/qwen2_5_vl/siglip2) leave
  `model.children` empty and encode the tree only in dotted `path_pattern`s. The
  viewer detects orphaned components/repeats and reattaches them by path,
  synthesizing intermediate containers (`vision_tower`, `language_model`, …) so
  every model is spreadable. Well-linked models are untouched.
- You can also **upload** an IR JSON file or **paste** artifact JSON directly.
- You can **paste / edit a config** (the merged referenced + salient fields) and
  click *Apply & resolve* to re-resolve all symbolic repeat counts live.
- Or pick a **Hub checkpoint** from the dropdown to fetch its `config.json`
  (both the Hub models API and `resolve/main/config.json` are called directly
  from the browser β€” no backend; they return permissive CORS headers). Gated or
  private checkpoints report a clear message instead of loading.

### Where to export from the generator

Write straight to **`ir/`** (its current output dir):

```
ir/manifest.json           # architecture index
ir/artifacts/<model>.json  # one artifact per architecture
```

The `IR_BASE` constant at the top of `js/app.js` controls this location; change
it there if you want the app to read from somewhere else.

## Project layout

```
index.html          # app shell: header, sidebar, canvas, inspector
style.css           # Hugging Face-flavoured styling
js/ir.js            # IR model + client-side repeat-count resolver
js/layout.js        # nested layered graph layout derived from the IR
js/graph.js         # SVG renderer + pan/zoom/select/collapse
js/app.js           # wiring: loading, sidebar controls, inspector (IR_BASE here)
ir/manifest.json    # architecture index β€” the generator writes here
ir/artifacts/*.json # IR artifacts β€” the app reads these directly
```

## Controls

- **Scroll** to zoom, **drag the canvas** to pan, **Fit** to recenter.
- **Drag a node** to reposition it (dragging a group box moves its whole
  subtree); edges follow live. **Reset layout** clears manual positions.
- **◐** toggles light / dark theme (defaults to the HF dark palette).
- **Click** a node to inspect it; click an edge entry in the inspector to jump.
- **βŠ• / βŠ–** on a repeat block expands / collapses it (repeats are **expanded by
  default**). *Collapse / Expand repeats* toggle all at once.
- Sidebar: architecture selector, Hub-checkpoint dropdown, config editor,
  edge-kind toggles, config-agnostic and provenance options.

## Known limitations (prototype)

- Layout is a lightweight custom layered algorithm; very large/expanded graphs
  are readable but not crossing-optimal (no ELK/Sugiyama ordering pass).
- Expanded repeats render one representative body instance (plus a stacked-deck
  hint in config-agnostic mode), not N unrolled copies β€” matching the IR.
- Edge routing is straight bezier (no orthogonal routing / obstacle avoidance).
- Non-repeat containers are always expanded; only symbolic repeats collapse.
- The Hub dropdown filters by the `model_type` tag, top 30 by downloads β€” not an
  exhaustive search; gated checkpoints can't expose their config to the browser.
- No search / filter across components yet.

## Next steps

- Optional ELK.js layout backend for large expanded graphs.
- Component search + "focus subtree" navigation.
- Deep-link state (selected node, expand set, config) in the URL.
- Hub integration: load the IR artifact itself by `model_id` (today only the
  `config.json` is fetched from the Hub; the IR still comes from `ir/`).
- Diff view between two configs (e.g. base vs. resolved layer counts).