Skip to content

Dropout: A Simple Way to Prevent Neural Networks from Overfitting

Authors: Geoffrey E. Hinton, Nitish Srivastava, Alex Krizhevsky, Ilya Sutskever, Ruslan R. Salakhutdinov Year: 2012 Venue: JMLR 2012 Citations: 40,000+

Summary

Introduces Dropout, a simple regularization technique that randomly drops units during training. Prevents co-adaptation and improves generalization. Simple, effective, widely used technique.

Key Concepts

  • Random Unit Dropping: Randomly zero activations during forward pass
  • Preventing Co-adaptation: Reduces dependence on specific neurons
  • Ensemble Effect: Equivalent to training many thinned networks
  • Inverted Dropout: Scales activations at train time, not test time
  • Regularization: Effective overfitting prevention
  • Model Averaging: Testing averages predictions of dropped networks

Impact

  • 40,000+ citations
  • Simple and effective regularization technique
  • Still widely used in modern networks
  • Inspired other stochastic regularization methods
  • Easy to implement and understand
  • Good baseline regularization method
  • Influenced understanding of regularization via noise

Key Results

  • Reduced overfitting on diverse datasets
  • Improved generalization significantly
  • Simple implementation enabling wide adoption
  • Effective alternative to other regularization
  • DropConnect (Wan et al., 2013)
  • Variational Dropout (Gal & Ghahramani, 2016)
  • Batch Normalization (Ioffe & Szegedy, 2015)
  • L1/L2 Regularization (Classical machine learning)