JEV-as-a-Judge: Accept When Confident, Escalate When Unsure Paper • 2609.26550 • Published 4 days ago • 32
MoME: Mixture-of-Memory Embeddings for Context-Aware Sparse Lookup Paper • 2609.15126 • Published 12 days ago • 12
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 9 days ago • 54
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Paper • 2609.20744 • Published 9 days ago • 51 • 4
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Paper • 2609.20744 • Published 9 days ago • 51
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 9 days ago • 180
JEPA-Anything: Learning Predictive Models across Different Worlds Paper • 2609.20800 • Published 9 days ago • 72
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 12 days ago • 50
Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 18 days ago • 81
view article Article Profiling in PyTorch (Part 3): Attention is all you profile +2 ariG23498, sergiopaniego, sayakpaul, ror • Jul 10 • 50
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 16 days ago • 700
The Price of Sparsity: Sufficient Conditions for Sparse Recovery using Sparse and Sparsified Measurements Paper • 2509.01809 • Published 18 days ago • 4
view article Article Bringing Nunchaku 4-bit Diffusion Inference to Diffusers rootonchair, sayakpaul • Jul 23 • 69
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 23 days ago • 113