Transformer Architectures (2017-2023)¶
The foundational architecture and its evolution for modern NLP.
Core Transformer¶
- Attention Is All You Need - The foundational transformer paper
BERT & Variants¶
- BERT: Pre-training of Deep Bidirectional Transformers - Masked language modeling
- What Does BERT Learn About the Structure of Language? - BERT analysis
GPT & Decoder Models¶
- Language Models are Unsupervised Multitask Learners (GPT-2) - Autoregressive LM
- Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer (T5) - Unified framework
Position Embeddings¶
- RoFormer: Enhanced Transformer with Rotary Position Embedding - Rotary position embeddings (RoPE)
- ALiBi: Train Short, Test Long - Attention with linear biases
Efficient Attention¶
- Introducing Sparse Attention - Sparse attention patterns
- FlashAttention: Fast and Memory-Efficient Exact Attention - IO-aware attention
- FlashAttention-2: Faster Attention with Better Parallelism - Further optimization
Attention Visualization¶
- What Does BERT Learn About the Structure of Language? - Understanding BERT
- Visualizing Attention in Transformer-Based Language Representation Models - Attention analysis
Key Insight: Transformers revolutionized NLP by enabling parallel processing and scaled to build all modern LLMs.