HYPIC: Accelerating Hybrid-Attention LLM Serving with Position-Independent Caching Paper • 2607.01299 • Published Jul 12 • 3
SmartThinker: Progressive Chain-of-Thought Length Calibration for Efficient Large Language Model Reasoning Paper • 2603.08000 • Published May 31 • 2
ExFold: Unified Expert Folding for Training-Free MoE Prefill-Decode Acceleration Paper • 2608.24938 • Published 6 days ago • 3
AgentOPSD: Recursive Self-Distillation for Agentic Reinforcement Learning Paper • 2608.05987 • Published 24 days ago • 100
Meta Llama 3 Collection This collection hosts the transformers and original repos of the Meta Llama 3 and Llama Guard 2 releases • 5 items • Updated Dec 6, 2024 • 997