title: Transformers Architecture Inspector
emoji: π
colorFrom: gray
colorTo: yellow
sdk: static
app_file: index.html
pinned: false
short_description: Inspect the semantic architecture IR of Transformers models
Transformers Architecture Viewer
An interactive, web-based viewer for the Transformers Architecture IR β a compact, semantic, config-parametric representation of a model architecture.
The viewer consumes the IR artifacts directly (it never inspects Transformers Python code, never loads checkpoint weights, and never instantiates a model). The graph layout is derived entirely from the IR: semantic components become nodes, symbolic repeats become collapsible blocks, and coarse dataflow becomes typed edges.
This is a UI prototype for a standalone Hugging Face Space, not the final production Hub integration.
What it shows
- Semantic components as nodes, colour-coded by kind (attention, feed-forward, normalization, embedding, position, β¦).
- Symbolic repeats as collapsible blocks, labelled e.g.
Decoder Layer Γ 32(resolved against a config) or, in config-agnostic mode,Decoder Layer Γ num_hidden_layersβ expanding a repeat then draws a stacked "deck" with aβ― repeated β―hint to convey the fully-deployed model without unrolling all N instances. - Hub checkpoints: a dropdown lists the most-downloaded public models on the
Hugging Face Hub that use this architecture; picking one loads its real
config.jsonand re-resolves every repeat count live. - Hierarchy as containment boxes and coarse dataflow as arrows, rendered separately.
- Edge kinds as visually distinct arrows:
data,residual,mask,position,cross_attentionβ each toggleable. - Provenance, per-component attributes (semantic facts), model-level architecture facts, and config-derived values for the selected component.
- On-node captions: each node shows a shape/attribute line derived from the
IR β e.g.
[B, S, 4096](observed dataflow, config-resolved) orhead dim 128 Β· n heads 32(attributes). Toggle with Show shapes & attributes. - Colour-by-role node fills (attention / feed-forward / normalization / β¦).
- Manual repositioning: drag any node/group to declutter edges.
Comparison mode
The β Compare button opens a two-architecture diff. It picks a comparison regime from the data and renders accordingly:
- Step 0 β regime: if one model
extendsthe other (ormodularitynames a parent) β same lineage (align by stable id). Otherwise it measures stable-id overlap: high β shared ids (align by id); low β cross-lineage (align by semantic kind lanes). Amodular_graph.jsonforest file would let this also suggest lineage comparisons β the code reads lineage from each artifact'sextends/modularitytoday. - Step 1 β headline: the
architecturefacts side-by-side (decoder Β· MHA Β· ropevsenc_dec Β· MHA Β· relative), diffs highlighted. - Step 3 β structure: matched nodes'
attributesdiffed (e.g. GQA shows up asn_kv_heads 32 β 8); added / removed nodes flagged. Cross-lineage pairs fall back to kind lanes (counts + representative attribute deltas per kind). - Step 4 β scale: hidden size, resolved depth (
repeatsΓconfig), intermediate size, vocab. Pick a Hub checkpoint per side to make it concrete. - Step 5 β topology: edge-kind presence β surfaces
cross_attention(enc-dec) andcache_*(decoder KV cache) that attribute diffs miss.
Side-by-side graphs: both architectures render in synced panels (pan/zoom move together). Components present in both models β matched by stable id β get a green "shared" outline in each panel, so common structure pops out while the differences stand alone. Below the graphs sits the full textual diff report.
The viewer owns all alignment + delta math; there is no shipped pairwise-diff. Same-lineage overlay rendering (ghosted base + highlighted patches on a single graph) is the next step; today the side-by-side view already highlights shared components. Current samples (llama/bert/t5) are standalone and cross-family, so they share few components β the shared outline lights up for lineage pairs (e.g. gemma vs llama) once those artifacts exist.
How close can it get to a hand-drawn architecture poster?
The visual language (nested dashed containers, class-name headers, Γ N
badges, colour-by-role, shape/attribute captions, config-resolved dims) is fully
IR-driven and rendered generically. Two things in a richly hand-drawn diagram
are not reachable from the current IR without either deepening the generator
or hardcoding one architecture:
- Sub-module detail β
q/k/v/o_proj,gate/up/down_proj, the SDPA box:LlamaAttention/LlamaMLPare leaves (children: []) in the IR. If the generator descends into them (and annotates dims/attributes), the viewer shows them automatically β no viewer change needed. - Task-head / runtime extras β LM head, softmax, logits, tokenizer chips,
the causal-mask heatmap: the artifact is the base model (
LlamaModel), and tokens/masks are runtime, not architecture. These would need new IR content (e.g. aβ¦ForCausalLMartifact) or a separate data source.
IR spec
The viewer conforms to the current architecture-template-v0 artifacts (as
emitted by the generator into ir/). It does not carry any legacy shims β
regenerate the artifacts and the viewer follows.
- Config block: reads
config.referenced_fields(authoritative parametric surface) +config.salient_fields(curated scalars), merged (referenced wins) to drive the resolver and the editable config. The fullconfig.fieldsblob is intentionally no longer produced or read. - Edge kinds are discovered from the artifact, so
cache_read/cache_write(and any future kind) get a colour, an arrowhead and a legend toggle. Kinds not in the known set fall back to a deterministic palette colour. - Pseudo endpoints:
input:*(e.g. attention mask) andstate:*(e.g. the KV cache) render as distinct source/sink pills so cache-read/write arrows have visible endpoints. - Semantic facts: model-level
architecture(family / view / positional / attention variant / MoE) shows as a sidebar summary and on the model node; per-componentattributesand themoekind appear in the inspector. - Observed dataflow: the
dataflowblock's per-stage symbolized shapes (e.g.[B, S, config.hidden_size]) are shown in the inspector β whole-model input/output on the model node, and in/out shapes on each captured component. - Modular inheritance: when
modularity.is_modularis set, the sidebar shows the parent model, patch count anddiff_size.
Tech choices
Zero-build, dependency-free static site (vanilla JS + SVG):
- No npm install, no bundler, no CDN β so it is CSP-safe and trivial to host on a
sdk: staticSpace. - Runs by serving the folder with any static file server.
- The layout engine (
js/layout.js) is a small nested/layered graph layout; the resolver (js/ir.js) is a tiny client-side evaluator for repeat count expressions such asconfig.num_hidden_layers.
Run locally
Any static server works (a server is needed because the app uses fetch +
ES modules, which browsers block on file://):
# from the repository root
python3 -m http.server 8000
# then open http://localhost:8000
or
npx serve . # http://localhost:3000
Deploy as a Hugging Face Space
One command pushes the app + everything under ir/ to a static Space
(creates it on first run, then re-uploads on each call):
./deploy.sh # β lysandre/transformers-architecture-inspector
./deploy.sh <owner>/<space-name> # β your own Space
Requires the hf CLI (pip install -U huggingface_hub) and a logged-in session
(hf auth login). The script runs hf repo create --type space --space-sdk static
then hf upload . . (excluding .git, .idea, deploy.sh). The Space serves
index.html at the root and the app fetches ir/manifest.json + its artifacts.
How it loads IR artifacts
The app reads the generator's output directly β ir/ is the single source of
truth, there is no copy step. Point your generator's export at ir/ and a
reload picks up the new files.
- On load it fetches
ir/manifest.jsonand populates the architecture selector frommanifest.architectures[](currentlyllama,bert,t5). - Manifest artifact paths (e.g.
artifacts/llama.json) are resolved relative to the manifest, i.e.ir/artifacts/llama.json. - Hierarchy reconstruction: most artifacts link submodules via
children, but some (multimodal models like gemma3/internvl/qwen2_5_vl/siglip2) leavemodel.childrenempty and encode the tree only in dottedpath_patterns. The viewer detects orphaned components/repeats and reattaches them by path, synthesizing intermediate containers (vision_tower,language_model, β¦) so every model is spreadable. Well-linked models are untouched. - You can also upload an IR JSON file or paste artifact JSON directly.
- You can paste / edit a config (the merged referenced + salient fields) and click Apply & resolve to re-resolve all symbolic repeat counts live.
- Or pick a Hub checkpoint from the dropdown to fetch its
config.json(both the Hub models API andresolve/main/config.jsonare called directly from the browser β no backend; they return permissive CORS headers). Gated or private checkpoints report a clear message instead of loading.
Where to export from the generator
Write straight to ir/ (its current output dir):
ir/manifest.json # architecture index
ir/artifacts/<model>.json # one artifact per architecture
The IR_BASE constant at the top of js/app.js controls this location; change
it there if you want the app to read from somewhere else.
Project layout
index.html # app shell: header, sidebar, canvas, inspector
style.css # Hugging Face-flavoured styling
js/ir.js # IR model + client-side repeat-count resolver
js/layout.js # nested layered graph layout derived from the IR
js/graph.js # SVG renderer + pan/zoom/select/collapse
js/app.js # wiring: loading, sidebar controls, inspector (IR_BASE here)
ir/manifest.json # architecture index β the generator writes here
ir/artifacts/*.json # IR artifacts β the app reads these directly
Controls
- Scroll to zoom, drag the canvas to pan, Fit to recenter.
- Drag a node to reposition it (dragging a group box moves its whole subtree); edges follow live. Reset layout clears manual positions.
- β toggles light / dark theme (defaults to the HF dark palette).
- Click a node to inspect it; click an edge entry in the inspector to jump.
- β / β on a repeat block expands / collapses it (repeats are expanded by default). Collapse / Expand repeats toggle all at once.
- Sidebar: architecture selector, Hub-checkpoint dropdown, config editor, edge-kind toggles, config-agnostic and provenance options.
Known limitations (prototype)
- Layout is a lightweight custom layered algorithm; very large/expanded graphs are readable but not crossing-optimal (no ELK/Sugiyama ordering pass).
- Expanded repeats render one representative body instance (plus a stacked-deck hint in config-agnostic mode), not N unrolled copies β matching the IR.
- Edge routing is straight bezier (no orthogonal routing / obstacle avoidance).
- Non-repeat containers are always expanded; only symbolic repeats collapse.
- The Hub dropdown filters by the
model_typetag, top 30 by downloads β not an exhaustive search; gated checkpoints can't expose their config to the browser. - No search / filter across components yet.
Next steps
- Optional ELK.js layout backend for large expanded graphs.
- Component search + "focus subtree" navigation.
- Deep-link state (selected node, expand set, config) in the URL.
- Hub integration: load the IR artifact itself by
model_id(today only theconfig.jsonis fetched from the Hub; the IR still comes fromir/). - Diff view between two configs (e.g. base vs. resolved layer counts).