-
Extending LLMs' Context Window with 100 Samples
Paper • 2401.07004 • Published • 16 -
Extending Context Window of Large Language Models via Semantic Compression
Paper • 2312.09571 • Published • 16 -
RULER: What's the Real Context Size of Your Long-Context Language Models?
Paper • 2404.06654 • Published • 42
Collections
Discover the best community collections!
Collections trending this week
-
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
Paper • 2401.10032 • Published • 13 -
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
Paper • 2401.04658 • Published • 27 -
FreeInit: Bridging Initialization Gap in Video Diffusion Models
Paper • 2312.07537 • Published • 27 -
TCNCA: Temporal Convolution Network with Chunked Attention for Scalable Sequence Processing
Paper • 2312.05605 • Published • 4
-
Analyzing and Improving the Training Dynamics of Diffusion Models
Paper • 2312.02696 • Published • 33 -
Adaptive Guidance: Training-free Acceleration of Conditional Diffusion Models
Paper • 2312.12487 • Published • 9 -
Implicit Diffusion: Efficient Optimization through Stochastic Sampling
Paper • 2402.05468 • Published • 6
-
Extending LLMs' Context Window with 100 Samples
Paper • 2401.07004 • Published • 16 -
Extending Context Window of Large Language Models via Semantic Compression
Paper • 2312.09571 • Published • 16 -
RULER: What's the Real Context Size of Your Long-Context Language Models?
Paper • 2404.06654 • Published • 42
-
FreGrad: Lightweight and Fast Frequency-aware Diffusion Vocoder
Paper • 2401.10032 • Published • 13 -
Lightning Attention-2: A Free Lunch for Handling Unlimited Sequence Lengths in Large Language Models
Paper • 2401.04658 • Published • 27 -
FreeInit: Bridging Initialization Gap in Video Diffusion Models
Paper • 2312.07537 • Published • 27 -
TCNCA: Temporal Convolution Network with Chunked Attention for Scalable Sequence Processing
Paper • 2312.05605 • Published • 4
-
Analyzing and Improving the Training Dynamics of Diffusion Models
Paper • 2312.02696 • Published • 33 -
Adaptive Guidance: Training-free Acceleration of Conditional Diffusion Models
Paper • 2312.12487 • Published • 9 -
Implicit Diffusion: Efficient Optimization through Stochastic Sampling
Paper • 2402.05468 • Published • 6