Skip to content

Attention Is All You Need

Authors: Vaswani et al. Year: 2017 ArXiv/Link: https://arxiv.org/abs/1706.03762

Summary

The foundational transformer architecture that introduced multi-head self-attention mechanisms, enabling parallel processing and better capture of long-range dependencies.

Key Concepts

  • Self-attention
  • Positional encoding
  • Multi-head attention
  • Encoder-decoder architecture
  • Scaled dot-product attention

Impact

Foundational architecture for all modern language models and transformers

Category

Transformers