Official Transcoder Release: Cross-Layer Sparse Autoencoders for Frontier Agentic LLMs

Paper HuggingFace Models HuggingFace Dataset License: Apache 2.0

Official release repository containing production-grade MLP Transcoders (cross-layer Sparse Autoencoders) for the NeurIPS 2026 paper:
"How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression"


πŸ“Œ Overview

Transcoders replace standard MLP sublayers by learning sparse, interpretable feature representations across layer transitions: a=Activation(xWencT+benc),y^=aWdec+bdec\mathbf{a} = \mathrm{Activation}(\mathbf{x} \mathbf{W}_{\mathrm{enc}}^T + \mathbf{b}_{\mathrm{enc}}), \quad \mathbf{\hat{y}} = \mathbf{a} \mathbf{W}_{\mathrm{dec}} + \mathbf{b}_{\mathrm{dec}} where $\mathbf{x} \in \mathbb{R}^{d_{\mathrm{model}}}$ is the input activation entering the target MLP, and $\mathbf{\hat{y}} \in \mathbb{R}^{d_{\mathrm{model}}}$ reconstructs the MLP output $\mathbf{y}$.

By projecting dense MLP updates into high-dimensional overcomplete feature dictionaries ($d_{\mathrm{feature}} = 102,400 \sim 204,800$, expansion factor $25\times \sim 40\times$), these checkpoints allow researchers to:

  1. Decompose complex internal computations into discrete, interpretable linear directions.
  2. Probe causal feature attribution and suppression mechanisms governing action initiation.
  3. Perform bidirectional feature steering without destabilizing native language capabilities.

πŸ“Š 1. Transcoder Performance & Metrics

All released Transcoder checkpoints were evaluated on held-out standard language corpus (EleutherAI/SmolLM2-135M-10B) to verify general reconstruction fidelity and sparsity. Reconstruction quality is measured via Variance Explained ($\mathrm{VE}$): VE=1βˆ’Var(yβˆ’y^)Var(y)\mathrm{VE} = 1 - \frac{\mathrm{Var}(\mathbf{y} - \mathbf{\hat{y}})}{\mathrm{Var}(\mathbf{y})} alongside the average number of active features per token ($\text{每 Token } L_0$).

Performance Summary Across Released Checkpoints

Model / ζ¨‘εž‹ Layer / 层号 SmolLM ζ–‡ζœ¬ VE 每 Token $L_0$
Mistral-Small-3.2-24B L22 59.44% 150.0
L23 63.34% 149.9
L24 65.31% 149.9
L25 65.14% 149.8
Granite-3.3-8B-Instruct L29 46.91% 250.0
L30 48.26% 250.0
L31 65.68% 250.0
L32 68.16% 250.0
L33 73.08% 250.0
L34 76.64% 250.0
Qwen3.5-9B L25 42.85% 75.8
L26 68.52% 137.5
L27 63.93% 138.5
L28 62.04% 165.0
Qwen3.5-4B L28 70.22% 250.0
L29 75.20% 250.0
L30 77.65% 250.0
L31 89.22% 250.0

Summary: Across all 4 frontier model families, the Transcoders deliver strong reconstruction fidelity on standard text (VE reaching up to 89.22%) while adhering to strictly bounded sparsity budgets ($L_0 \in [75, 250]$), confirming that the learned dictionaries capture genuine computational transformations rather than dense approximations.


πŸ’‘ 2. Practical Experimental Recommendations

Based on our empirical methodology and cross-architecture validations, we offer two primary recommendations for researchers training, fine-tuning, or applying Transcoders to downstream mechanistic interpretability tasks:

β‘  Inject Domain-Relevant Text into Training to Prevent Out-Of-Distribution (OOD) Collapse

  • Observation: Training Transcoders exclusively on raw, open-domain text distributions makes them susceptible to severe out-of-distribution (OOD) collapse when later evaluated on specialized prompts (such as agentic prompts with complex system framing, structured tool declarations, and schema syntax). On such prompts, encoder activations can shift beyond nominal ranges, leading to uncalibrated feature responses and dead latents.
  • Recommendation: In training or downstream adaptation mixtures, strongly recommend injecting a targeted proportion (e.g., $0.05% \sim 1.0%$) of domain-relevant text (such as agentic interactions, tool specifications, or structured workflows). Crucially, this mixed text should be conceptually aligned with the target behavior but strictly disjoint from any held-out evaluation or probing prompts to ensure methodological validity. For specialized downstream applications, practitioners are similarly encouraged to fine-tune or adapt the Transcoder on target-relevant distributions using this mixing strategy to prevent OOD failure modes.

β‘‘ Employ Empirically Calibrated Initialization & FP32 Optimizer Precision

  • Calibrated Initialization: Uncalibrated or naive zero initializations can cause large initial reconstruction errors and early feature atrophy. Pre-computing empirical mean input and output activations over calibration tokens ($\bar{\mathbf{x}} = \mathbb{E}[\mathbf{x}]$, $\bar{\mathbf{y}} = \mathbb{E}[\mathbf{y}]$) to calibrate biases: $$\mathbf{b}{\mathrm{dec}} \leftarrow \bar{\mathbf{y}}, \quad \mathbf{b}{\mathrm{enc}} \leftarrow -\bar{\mathbf{x}} \mathbf{W}_{\mathrm{enc}}^T$$ ensures the model starts from the optimal baseline mean reconstruction and that encoder inputs are properly centered around activation thresholds.
  • FP32 Optimizer Precision: Given the high dimensionality of Transcoder parameter matrices and the extreme sparsity of feature activations ($L_0 \ll d_{\mathrm{feature}}$), individual features receive infrequent and localized gradient signals. Using FP32 precision for optimizer states (such as AdamW momentum and variance accumulators) has proven essential. It prevents sparse updates from vanishing due to floating-point truncation, ensuring stable, monotonic convergence across all model depths.

πŸ“‚ Released Checkpoint Structure

release/
β”œβ”€β”€ granite-3.3-8b-instruct/
β”‚   β”œβ”€β”€ release_manifest.json
β”‚   β”œβ”€β”€ transcoder_layer29.pt      # TopK-250
β”‚   β”œβ”€β”€ transcoder_layer30.pt      # TopK-250
β”‚   β”œβ”€β”€ transcoder_layer31.pt      # TopK-250
β”‚   β”œβ”€β”€ transcoder_layer32.pt      # TopK-250
β”‚   β”œβ”€β”€ transcoder_layer33.pt      # TopK-250
β”‚   β”œβ”€β”€ transcoder_layer34.pt      # TopK-250
β”‚   └── transcoder_layer35.pt      # TopK-250
β”œβ”€β”€ mistral-small-3.2-24b/
β”‚   β”œβ”€β”€ release_manifest.json
β”‚   β”œβ”€β”€ transcoder_layer22.pt      # TopK-150
β”‚   β”œβ”€β”€ transcoder_layer23.pt      # TopK-150
β”‚   β”œβ”€β”€ transcoder_layer24.pt      # TopK-150
β”‚   └── transcoder_layer25.pt      # TopK-150
β”œβ”€β”€ qwen35-4b/
β”‚   β”œβ”€β”€ release_manifest.json
β”‚   β”œβ”€β”€ transcoder_layer28.pt      # TopK-250
β”‚   β”œβ”€β”€ transcoder_layer29.pt      # TopK-250
β”‚   β”œβ”€β”€ transcoder_layer30.pt      # TopK-250
β”‚   └── transcoder_layer31.pt      # TopK-250
└── qwen35-9b/
    β”œβ”€β”€ release_manifest.json
    β”œβ”€β”€ transcoder_layer25.pt      # ReLU (L1/tanh sparsity)
    β”œβ”€β”€ transcoder_layer26.pt      # ReLU (L1/tanh sparsity)
    β”œβ”€β”€ transcoder_layer27.pt      # ReLU (L1/tanh sparsity)
    β”œβ”€β”€ transcoder_layer28.pt      # ReLU (L1/tanh sparsity)
    β”œβ”€β”€ transcoder_layer29.pt      # TopK-250
    └── transcoder_layer30.pt      # TopK-250

πŸš€ Quick Loading Example

import torch
from huggingface_hub import hf_hub_download

# Download checkpoint from Hugging Face Hub (or use local path)
checkpoint_path = hf_hub_download(
    repo_id="XijieGong/MI4ToolCalling",
    filename="mistral-small-3.2-24b/transcoder_layer25.pt"
)
payload = torch.load(checkpoint_path, map_location="cpu")

# Extract state dictionary
state_dict = payload.get("model", payload)
W_enc = state_dict["W_enc"]  # Shape: (d_feature, d_model)
b_enc = state_dict["b_enc"]  # Shape: (d_feature,)
W_dec = state_dict["W_dec"]  # Shape: (d_feature, d_model)
b_dec = state_dict["b_dec"]  # Shape: (d_model,)

print(f"Loaded Transcoder: {W_enc.shape[0]} features, d_model={W_enc.shape[1]}")

πŸ“– Citation

@inproceedings{mi4toolcalling2026,
  title={How Do Agentic LLMs Decide to Call Tools? A Tool-Call Vector Shaped by Suppression},
  author={Anonymous Authors},
  booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
  year={2026}
}

πŸ“œ License

This repository and all released Transcoder checkpoint weights are licensed under the Apache License 2.0.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support