Attention Is All You Need¶
Authors: Vaswani et al. Year: 2017 ArXiv/Link: https://arxiv.org/abs/1706.03762
Summary¶
The foundational transformer architecture that introduced multi-head self-attention mechanisms, enabling parallel processing and better capture of long-range dependencies.
Key Concepts¶
- Self-attention
- Positional encoding
- Multi-head attention
- Encoder-decoder architecture
- Scaled dot-product attention
Impact¶
Foundational architecture for all modern language models and transformers
Category¶
Transformers