Collections
Discover the best community collections!
Collections trending this week
-
QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models
Paper • 2310.16795 • Published • 27 -
Pareto-Optimal Quantized ResNet Is Mostly 4-bit
Paper • 2105.03536 • Published • 3 -
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
Paper • 2403.15447 • Published • 15
-
Adding NVMe SSDs to Enable and Accelerate 100B Model Fine-tuning on a Single GPU
Paper • 2403.06504 • Published • 56 -
Token-Level Adaptation of LoRA Adapters for Downstream Task Generalization
Paper • 2311.10847 • Published • 2 -
PERL: Parameter Efficient Reinforcement Learning from Human Feedback
Paper • 2403.10704 • Published • 60
-
Multistep Consistency Models
Paper • 2403.06807 • Published • 15 -
Improving Text-to-Image Consistency via Automatic Prompt Optimization
Paper • 2403.17804 • Published • 19 -
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Paper • 2404.01197 • Published • 31 -
Consistency Flow Matching: Defining Straight Flows with Velocity Consistency
Paper • 2407.02398 • Published • 18
-
SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts
Paper • 2105.03036 • Published • 2 -
Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition
Paper • 2112.05820 • Published • 2 -
SpeechMoE2: Mixture-of-Experts Model with Improved Routing
Paper • 2111.11831 • Published • 2
-
Multistep Consistency Models
Paper • 2403.06807 • Published • 15 -
Improving Text-to-Image Consistency via Automatic Prompt Optimization
Paper • 2403.17804 • Published • 19 -
Getting it Right: Improving Spatial Consistency in Text-to-Image Models
Paper • 2404.01197 • Published • 31 -
Consistency Flow Matching: Defining Straight Flows with Velocity Consistency
Paper • 2407.02398 • Published • 18
-
QMoE: Practical Sub-1-Bit Compression of Trillion-Parameter Models
Paper • 2310.16795 • Published • 27 -
Pareto-Optimal Quantized ResNet Is Mostly 4-bit
Paper • 2105.03536 • Published • 3 -
Decoding Compressed Trust: Scrutinizing the Trustworthiness of Efficient LLMs Under Compression
Paper • 2403.15447 • Published • 15
-
SpeechMoE: Scaling to Large Acoustic Models with Dynamic Routing Mixture of Experts
Paper • 2105.03036 • Published • 2 -
Building a great multi-lingual teacher with sparsely-gated mixture of experts for speech recognition
Paper • 2112.05820 • Published • 2 -
SpeechMoE2: Mixture-of-Experts Model with Improved Routing
Paper • 2111.11831 • Published • 2
-
Adding NVMe SSDs to Enable and Accelerate 100B Model Fine-tuning on a Single GPU
Paper • 2403.06504 • Published • 56 -
Token-Level Adaptation of LoRA Adapters for Downstream Task Generalization
Paper • 2311.10847 • Published • 2 -
PERL: Parameter Efficient Reinforcement Learning from Human Feedback
Paper • 2403.10704 • Published • 60