Skip to content

Rectified Linear Units Improve Restricted Boltzmann Machines

Authors: Nair, V., & Hinton, G. E. Year: 2010 ArXiv: https://proceedings.mlr.press/v15/nair10a.html

Summary

ReLU introduces the simple but effective max(0, x) activation function. Despite its simplicity, ReLU enabled training of much deeper networks than sigmoid or tanh, becoming the default activation function in modern deep learning.

Key Concepts

  • Non-saturation for positive values
  • Simple computation (max operation)
  • Sparsity-inducing properties
  • Enables training of very deep networks
  • Variants: Leaky ReLU, ELU, GELU

Impact

ReLU revolutionized deep learning by solving the vanishing gradient problem. It's used in nearly all modern neural networks and remains the most popular activation function after 15+ years.


Citation:

@inproceedings{nair2010rectified,
  title={Rectified linear units improve restricted boltzmann machines},
  author={Nair, Vinod and Hinton, Geoffrey E},
  booktitle={Proceedings of the 27th international conference on machine learning},
  pages={807--814},
  year={2010}
}