WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
Paper • 2609.23033 • Published • 9
Efficient AI
KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems
PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction