-
Attention Is All You Need
Paper • 1706.03762 • Published • 134 -
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Paper • 2205.14135 • Published • 15 -
Efficient Memory Management for Large Language Model Serving with PagedAttention
Paper • 2309.06180 • Published • 64 -
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Paper • 2307.08691 • Published • 9
Lin Cui
cuixx042
AI & ML interests
None yet
Recent Activity
updated a collection 15 days ago
AI updated a collection 15 days ago
AI updated a collection 17 days ago
AIOrganizations
None yet