Skip to content

FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness

Authors: Dao et al. Year: 2022 ArXiv/Link: https://arxiv.org/abs/2205.14135

Summary

Optimized attention computation through IO-aware algorithms, achieving 2x speedup and reduced memory usage without approximation.

Key Concepts

  • IO-awareness
  • Memory efficiency
  • GPU computation
  • Block-wise computation
  • Exact attention

Impact

Enabled practical training of very long context models

Category

Attention Mechanisms