KVCMAS: Efficient KV cache Correction for Shared Context in Multi-Agent Systems Paper • 2609.34060 • Published 3 days ago • 4
PReCache: Efficient KV Cache Sharing for Multi-LoRA Agents via Low-Rank Precomputation and Neutral Reconstruction Paper • 2609.34054 • Published 3 days ago • 4
CompactAttention: Accelerating Chunked Prefill with Block-Union KV Selection Paper • 2605.16839 • Published May 16 • 13
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents Paper • 2602.01053 • Published Feb 1 • 8
RelayGen: Intra-Generation Model Switching for Efficient Reasoning Paper • 2602.06454 • Published Feb 6 • 12
LRAgent: Efficient KV Cache Sharing for Multi-LoRA LLM Agents Paper • 2602.01053 • Published Feb 1 • 8
Token Sparse Attention: Efficient Long-Context Inference with Interleaved Token Selection Paper • 2602.03216 • Published Feb 3 • 14
LiteStage: Latency-aware Layer Skipping for Multi-stage Reasoning Paper • 2510.14211 • Published Oct 16, 2025 • 9
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models Paper • 2509.17428 • Published Sep 22, 2025 • 9
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models Paper • 2509.17428 • Published Sep 22, 2025 • 9
QWHA: Quantization-Aware Walsh-Hadamard Adaptation for Parameter-Efficient Fine-Tuning on Large Language Models Paper • 2509.17428 • Published Sep 22, 2025 • 9 • 2
L4Q: Parameter Efficient Quantization-Aware Training on Large Language Models via LoRA-wise LSQ Paper • 2402.04902 • Published Feb 7, 2024 • 5
Reasoning Path Compression: Compressing Generation Trajectories for Efficient LLM Reasoning Paper • 2505.13866 • Published May 20, 2025 • 17
AUTOACT: Automatic Agent Learning from Scratch via Self-Planning Paper • 2401.05268 • Published Jan 10, 2024 • 4
Running 4.05k The Ultra-Scale Playbook 🌌 4.05k The ultimate guide to training LLM on large GPU Clusters