Adam: A Method for Stochastic Optimization¶
Authors: Diederik P. Kingma, Jimmy Lei Ba Year: 2015 (arXiv 2014) Venue: ICLR 2015 Citations: 50,000+
Link¶
Summary¶
Introduces Adam (Adaptive Moment Estimation), an optimization algorithm combining adaptive learning rates with momentum. Computes adaptive per-parameter learning rates using gradient history. Becomes standard optimizer for deep learning.
Key Concepts¶
- Adaptive Learning Rates: Different learning rate per parameter
- Momentum: First moment (mean) of gradients
- Second Moment: Second moment (uncentered variance) of gradients
- Bias Correction: Corrects for initialization bias in moment estimates
- Efficient Computation: Simple to implement and compute
- Hyperparameter Robustness: Works well with default settings
Impact¶
- 50,000+ citations
- Most widely used optimizer in deep learning
- Default choice for training deep networks
- Inspired AdaBound, RAdam, and other variants
- Better generalization than SGD in many cases
- Practical and easy to implement
- Critical for modern deep learning
Key Results¶
- Faster convergence than SGD
- Better generalization on many tasks
- Robust across diverse problems
- Works with sparse gradients
Related Papers¶
- Adagrad: Adaptive Gradient Algorithm (Duchi et al., 2011)
- RMSProp (Tieleman & Hinton, 2012)
- Momentum and Nesterov (Sutskever et al., 2013)
- Convergence Analysis of Adam (Reddi et al., 2019)