GeoPair: Geometry-Preserving Cross-Layer Factorization for Training-Free Transformer Compression
Abstract
Transformer architectures exhibit cross-layer redundancies, yet post-training compression pipelines typically optimize layers in isolation or rely on heuristic grouping strategies that disregard layer-specific activation geometries. We introduce a principled, training-free framework that sequentially optimizes cross-layer weight pairings and shared-dictionary factorizations. Rather than forcing weights of adjacent layers to share a basis or heuristically merging activation statistics, our approach identifies structurally compatible projections and learns a shared representation that better preserves each layer's distinct calibration geometry. Coupled with structured sparsity, this yields highly efficient weight decompositions without sacrificing functional fidelity. Across diverse architectures, scales, and modalities, our method achieves state-of-the-art results, consistently outperforming independent structured weight decompositions and alternative pairwise weight factorizations, which operate under heuristic grouping strategies. By replacing heuristic engineering strategies with a convergent, optimization-driven pipeline, we establish a theoretically grounded foundation for scalable, transformer compression across different modalities.
Community
Solve w.r.t different spaces
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Mind the Approximation: Fisher-Weighted SVD Compression for ViTs (2026)
- LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry (2026)
- SHIFT-LLM: Distribution Shift Correction in Depth-Pruned LLMs (2026)
- REAL-Q: E2E LLM Quantization via Dynamic Gradient Descent (2026)
- Importance-Aware Low-Rank Distillation of Diffusion Transformers (2026)
- Correlation-Aware Structured Pruning for Large Language Models (2026)
- Per-Matrix Optimality Is Not Enough: Three-Level Optimization for Low-Rank LLM Compression (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.25963 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper