TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Paper • 2604.12012 • Published Apr 13 • 16
Conda: Column-Normalized Adam for Training Large Language Models Faster Paper • 2509.24218 • Published Sep 29, 2025 • 2
Leviathan: Decoupling Input and Output Representations in Language Models Paper • 2601.22040 • Published May 7 • 1
VersatileFFN: Achieving Parameter Efficiency in LLMs via Adaptive Wide-and-Deep Reuse Paper • 2512.14531 • Published Dec 16, 2025 • 16
TWEO: Transformers Without Extreme Outliers Enables FP8 Training And Quantization For Dummies Paper • 2511.23225 • Published Nov 28, 2025 • 3
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale Paper • 2601.22146 • Published Jan 29 • 12
Momentum Attention: The Physics of In-Context Learning and Spectral Forensics for Mechanistic Interpretability Paper • 2602.04902 • Published Feb 3 • 1