FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness¶
Authors: Dao et al. Year: 2022 ArXiv/Link: https://arxiv.org/abs/2205.14135
Summary¶
Optimized attention computation through IO-aware algorithms, achieving 2x speedup and reduced memory usage without approximation.
Key Concepts¶
- IO-awareness
- Memory efficiency
- GPU computation
- Block-wise computation
- Exact attention
Impact¶
Enabled practical training of very long context models
Category¶
Attention Mechanisms