HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching Paper • 2607.01299 • Published Jul 12 • 3
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration Paper • 2608.24938 • Published 5 days ago • 3
OSP-Next: Efficient High-Quality Video Generation with Sparse Sequence Parallelism, HiF8 Quantization, and Reinforcement Learning Paper • 2605.28691 • Published May 27 • 25
BEAM: Binary Expert Activation Masking for Dynamic Routing in MoE Paper • 2605.14438 • Published May 14 • 5
Rethinking Text-based Protein Understanding: Retrieval or LLM? Paper • 2505.20354 • Published May 26, 2025 • 1
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Paper • 2506.03147 • Published Jun 3, 2025 • 59