Official Transcoder Release: Cross-Layer Sparse Autoencoders for Frontier Agentic LLMs
Official release repository containing production-grade MLP Transcoders (cross-layer Sparse Autoencoders) for the NeurIPS 2026 paper:
"How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression"
π Overview
Transcoders replace standard MLP sublayers by learning sparse, interpretable feature representations across layer transitions: where $\mathbf{x} \in \mathbb{R}^{d_{\mathrm{model}}}$ is the input activation entering the target MLP, and $\mathbf{\hat{y}} \in \mathbb{R}^{d_{\mathrm{model}}}$ reconstructs the MLP output $\mathbf{y}$.
By projecting dense MLP updates into high-dimensional overcomplete feature dictionaries ($d_{\mathrm{feature}} = 102,400 \sim 204,800$, expansion factor $25\times \sim 40\times$), these checkpoints allow researchers to:
- Decompose complex internal computations into discrete, interpretable linear directions.
- Probe causal feature attribution and suppression mechanisms governing action initiation.
- Perform bidirectional feature steering without destabilizing native language capabilities.
π 1. Transcoder Performance & Metrics
All released Transcoder checkpoints were evaluated on held-out standard language corpus (EleutherAI/SmolLM2-135M-10B) to verify general reconstruction fidelity and sparsity. Reconstruction quality is measured via Variance Explained ($\mathrm{VE}$):
alongside the average number of active features per token ($\text{ζ― Token } L_0$).
Performance Summary Across Released Checkpoints
| Model / 樑ε | Layer / ε±ε· | SmolLM ζζ¬ VE | ζ― Token $L_0$ |
|---|---|---|---|
| Mistral-Small-3.2-24B | L22 | 59.44% | 150.0 |
| L23 | 63.34% | 149.9 | |
| L24 | 65.31% | 149.9 | |
| L25 | 65.14% | 149.8 | |
| Granite-3.3-8B-Instruct | L29 | 46.91% | 250.0 |
| L30 | 48.26% | 250.0 | |
| L31 | 65.68% | 250.0 | |
| L32 | 68.16% | 250.0 | |
| L33 | 73.08% | 250.0 | |
| L34 | 76.64% | 250.0 | |
| Qwen3.5-9B | L25 | 42.85% | 75.8 |
| L26 | 68.52% | 137.5 | |
| L27 | 63.93% | 138.5 | |
| L28 | 62.04% | 165.0 | |
| Qwen3.5-4B | L28 | 70.22% | 250.0 |
| L29 | 75.20% | 250.0 | |
| L30 | 77.65% | 250.0 | |
| L31 | 89.22% | 250.0 |
Summary: Across all 4 frontier model families, the Transcoders deliver strong reconstruction fidelity on standard text (VE reaching up to 89.22%) while adhering to strictly bounded sparsity budgets ($L_0 \in [75, 250]$), confirming that the learned dictionaries capture genuine computational transformations rather than dense approximations.
π‘ 2. Practical Experimental Recommendations
Based on our empirical methodology and cross-architecture validations, we offer two primary recommendations for researchers training, fine-tuning, or applying Transcoders to downstream mechanistic interpretability tasks:
β Inject Domain-Relevant Text into Training to Prevent Out-Of-Distribution (OOD) Collapse
- Observation: Training Transcoders exclusively on raw, open-domain text distributions makes them susceptible to severe out-of-distribution (OOD) collapse when later evaluated on specialized prompts (such as agentic prompts with complex system framing, structured tool declarations, and schema syntax). On such prompts, encoder activations can shift beyond nominal ranges, leading to uncalibrated feature responses and dead latents.
- Recommendation: In training or downstream adaptation mixtures, strongly recommend injecting a targeted proportion (e.g., $0.05% \sim 1.0%$) of domain-relevant text (such as agentic interactions, tool specifications, or structured workflows). Crucially, this mixed text should be conceptually aligned with the target behavior but strictly disjoint from any held-out evaluation or probing prompts to ensure methodological validity. For specialized downstream applications, practitioners are similarly encouraged to fine-tune or adapt the Transcoder on target-relevant distributions using this mixing strategy to prevent OOD failure modes.
β‘ Employ Empirically Calibrated Initialization & FP32 Optimizer Precision
- Calibrated Initialization: Uncalibrated or naive zero initializations can cause large initial reconstruction errors and early feature atrophy. Pre-computing empirical mean input and output activations over calibration tokens ($\bar{\mathbf{x}} = \mathbb{E}[\mathbf{x}]$, $\bar{\mathbf{y}} = \mathbb{E}[\mathbf{y}]$) to calibrate biases: $$\mathbf{b}{\mathrm{dec}} \leftarrow \bar{\mathbf{y}}, \quad \mathbf{b}{\mathrm{enc}} \leftarrow -\bar{\mathbf{x}} \mathbf{W}_{\mathrm{enc}}^T$$ ensures the model starts from the optimal baseline mean reconstruction and that encoder inputs are properly centered around activation thresholds.
- FP32 Optimizer Precision: Given the high dimensionality of Transcoder parameter matrices and the extreme sparsity of feature activations ($L_0 \ll d_{\mathrm{feature}}$), individual features receive infrequent and localized gradient signals. Using FP32 precision for optimizer states (such as AdamW momentum and variance accumulators) has proven essential. It prevents sparse updates from vanishing due to floating-point truncation, ensuring stable, monotonic convergence across all model depths.
π Released Checkpoint Structure
release/
βββ granite-3.3-8b-instruct/
β βββ release_manifest.json
β βββ transcoder_layer29.pt # TopK-250
β βββ transcoder_layer30.pt # TopK-250
β βββ transcoder_layer31.pt # TopK-250
β βββ transcoder_layer32.pt # TopK-250
β βββ transcoder_layer33.pt # TopK-250
β βββ transcoder_layer34.pt # TopK-250
β βββ transcoder_layer35.pt # TopK-250
βββ mistral-small-3.2-24b/
β βββ release_manifest.json
β βββ transcoder_layer22.pt # TopK-150
β βββ transcoder_layer23.pt # TopK-150
β βββ transcoder_layer24.pt # TopK-150
β βββ transcoder_layer25.pt # TopK-150
βββ qwen35-4b/
β βββ release_manifest.json
β βββ transcoder_layer28.pt # TopK-250
β βββ transcoder_layer29.pt # TopK-250
β βββ transcoder_layer30.pt # TopK-250
β βββ transcoder_layer31.pt # TopK-250
βββ qwen35-9b/
βββ release_manifest.json
βββ transcoder_layer25.pt # ReLU (L1/tanh sparsity)
βββ transcoder_layer26.pt # ReLU (L1/tanh sparsity)
βββ transcoder_layer27.pt # ReLU (L1/tanh sparsity)
βββ transcoder_layer28.pt # ReLU (L1/tanh sparsity)
βββ transcoder_layer29.pt # TopK-250
βββ transcoder_layer30.pt # TopK-250
π Quick Loading Example
import torch
from huggingface_hub import hf_hub_download
# Download checkpoint from Hugging Face Hub (or use local path)
checkpoint_path = hf_hub_download(
repo_id="XijieGong/MI4ToolCalling",
filename="mistral-small-3.2-24b/transcoder_layer25.pt"
)
payload = torch.load(checkpoint_path, map_location="cpu")
# Extract state dictionary
state_dict = payload.get("model", payload)
W_enc = state_dict["W_enc"] # Shape: (d_feature, d_model)
b_enc = state_dict["b_enc"] # Shape: (d_feature,)
W_dec = state_dict["W_dec"] # Shape: (d_feature, d_model)
b_dec = state_dict["b_dec"] # Shape: (d_model,)
print(f"Loaded Transcoder: {W_enc.shape[0]} features, d_model={W_enc.shape[1]}")
π Citation
@inproceedings{mi4toolcalling2026,
title={How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression},
author={Anonymous Authors},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2026}
}
π License
This repository and all released Transcoder checkpoint weights are licensed under the Apache License 2.0.