Introducing LLMs with Sparse Attention¶
Authors: Child et al. Year: 2019 ArXiv/Link: https://arxiv.org/abs/1904.10509
Summary¶
Sparse Transformers use structured sparsity in attention to enable longer sequences and more efficient computation.
Key Concepts¶
- Sparse attention
- Attention patterns
- Long sequences
- Computational efficiency
- Structured sparsity
Impact¶
Enabled longer context through efficient attention patterns
Category¶
Long Context & Scaling