Skip to content

Adam: A Method for Stochastic Optimization

Authors: Diederik P. Kingma, Jimmy Lei Ba Year: 2015 (arXiv 2014) Venue: ICLR 2015 Citations: 50,000+

Summary

Introduces Adam (Adaptive Moment Estimation), an optimization algorithm combining adaptive learning rates with momentum. Computes adaptive per-parameter learning rates using gradient history. Becomes standard optimizer for deep learning.

Key Concepts

  • Adaptive Learning Rates: Different learning rate per parameter
  • Momentum: First moment (mean) of gradients
  • Second Moment: Second moment (uncentered variance) of gradients
  • Bias Correction: Corrects for initialization bias in moment estimates
  • Efficient Computation: Simple to implement and compute
  • Hyperparameter Robustness: Works well with default settings

Impact

  • 50,000+ citations
  • Most widely used optimizer in deep learning
  • Default choice for training deep networks
  • Inspired AdaBound, RAdam, and other variants
  • Better generalization than SGD in many cases
  • Practical and easy to implement
  • Critical for modern deep learning

Key Results

  • Faster convergence than SGD
  • Better generalization on many tasks
  • Robust across diverse problems
  • Works with sparse gradients
  • Adagrad: Adaptive Gradient Algorithm (Duchi et al., 2011)
  • RMSProp (Tieleman & Hinton, 2012)
  • Momentum and Nesterov (Sutskever et al., 2013)
  • Convergence Analysis of Adam (Reddi et al., 2019)