Change the Product, Keep the Parameters: Associative Algebra Layers for Transformers
Abstract
Fast matrix multiplication algorithms keep the product fixed and search for a cheaper way to evaluate it. We instead ask whether a Transformer's learned projections can use a different, cheaper product altogether. Building on an associative-algebra construction that replaces ordinary matrix multiplication with a sparser interaction table over the same weight blocks, we construct a family with quadratic arithmetic in the matrix dimension when the physical block size remains fixed, and derive finite-shape constraints for GPU execution. The construction is provably optimal for its bilinear rank by the Alder--Strassen bound and can be realized as row-typed rectangular projections compatible with causal masking and KV-cached decoding. We provide an empirical test of this approach by training two approximately 110M-parameter decoder-only Transformer LMs from the same recipe and 12.3B-token budget, differing only in their feed-forward layer: one uses ordinary dense matrix multiplication and the other uses the associative-algebra product. Across four prompt domains, the algebraic model achieves a 6.2--7.8\% increase in end-to-end generation throughput, while obtaining lower scores on all three reported downstream metrics. We treat these results as a feasibility and trainability check for the proposed approach at small scale, leaving further investigation to future work.
Community
We study whether the matrix multiplication used in Transformer projections can be replaced by a different, cheaper associative product while keeping the same weight bank and parameter count.
We construct an associative algebra product with a lower bilinear rank than standard matrix multiplication. For the q=2 case, the product has rank 6, compared with rank 7 for Strassen's 2×2 algorithm. This allows us to replace linear projections without introducing additional nonlinearities, while reducing their arithmetic cost.
We evaluate the approach at three levels: GPU projection kernels across several Transformer shapes, including Qwen and DeepSeek; and two approximately 110M-parameter decoder-only LMs trained for 12.3B tokens, differing only in the MLP multiplication law. The algebraic model achieves a 6.2–7.8% increase in end-to-end generation throughput across four prompt domains, while obtaining lower scores on GSM8K, MBPP, and IFEval.
The results provide a small-scale feasibility and trainability study of changing the multiplication law of Transformer projections while retaining the full parameter bank.
Get this paper in your agent:
hf papers read 2609.32814 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper