Running 2 Sparsely gated tiny linear experts 🐥 A compute-efficient and interpretable transformer FFN layer