Collections
Discover the best community collections!
Collections trending this week
-
Large Language Models for Compiler Optimization
Paper • 2309.07062 • Published • 25 -
Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time
Paper • 2310.17157 • Published • 14 -
FP8-LM: Training FP8 Large Language Models
Paper • 2310.18313 • Published • 34 -
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
Paper • 2310.19102 • Published • 11
-
Large Language Models for Compiler Optimization
Paper • 2309.07062 • Published • 25 -
Deja Vu: Contextual Sparsity for Efficient LLMs at Inference Time
Paper • 2310.17157 • Published • 14 -
FP8-LM: Training FP8 Large Language Models
Paper • 2310.18313 • Published • 34 -
Atom: Low-bit Quantization for Efficient and Accurate LLM Serving
Paper • 2310.19102 • Published • 11