Skip to content

Introducing LLMs with Sparse Attention

Authors: Child et al. Year: 2019 ArXiv/Link: https://arxiv.org/abs/1904.10509

Summary

Sparse Transformers use structured sparsity in attention to enable longer sequences and more efficient computation.

Key Concepts

  • Sparse attention
  • Attention patterns
  • Long sequences
  • Computational efficiency
  • Structured sparsity

Impact

Enabled longer context through efficient attention patterns

Category

Long Context & Scaling