DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 9 days ago • 179
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 23 days ago • 113
Language Models Can Control Their Own Attention Paper • 2609.02737 • Published 24 days ago • 79
Gemma 4 Collection Our most intelligent open models to date • 16 items • Updated Aug 20 • 1.14k
REVE: A Foundation Model for EEG -- Adapting to Any Setup with Large-Scale Pretraining on 25,000 Subjects Paper • 2510.21585 • Published Oct 24, 2025 • 8
D5P4: Partition Determinantal Point Process for Diversity in Parallel Discrete Diffusion Decoding Paper • 2603.19146 • Published Jun 5
Inner Loop Inference for Pretrained Transformers: Unlocking Latent Capabilities Without Training Paper • 2602.14759 • Published Feb 16
Residual Connections and the Causal Shift: Uncovering a Structural Misalignment in Transformers Paper • 2602.14760 • Published Feb 16