Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Paper • 2608.12149 • Published 3 days ago • 18
ICA Lens: Interpreting Language Models Without Training Another Dictionary Paper • 2606.11722 • Published Jun 10 • 18
Hybrid Linear Attention Research Collection All 1.3B & 340M hybrid linear-attention experiments. • 62 items • Updated Sep 11, 2025 • 14