-
Priority Sampling of Large Language Models for Compilers
Paper • 2402.18734 • Published • 19 -
Accelerating Large Language Model Decoding with Speculative Sampling
Paper • 2302.01318 • Published • 5 -
Fast Inference from Transformers via Speculative Decoding
Paper • 2211.17192 • Published • 11 -
AttentiveNAS: Improving Neural Architecture Search via Attentive Sampling
Paper • 2011.09011 • Published • 2
Collections
Discover the best community collections!
Collections trending this week
-
Priority Sampling of Large Language Models for Compilers
Paper • 2402.18734 • Published • 19 -
Accelerating Large Language Model Decoding with Speculative Sampling
Paper • 2302.01318 • Published • 5 -
Fast Inference from Transformers via Speculative Decoding
Paper • 2211.17192 • Published • 11 -
AttentiveNAS: Improving Neural Architecture Search via Attentive Sampling
Paper • 2011.09011 • Published • 2