Rectified Linear Units Improve Restricted Boltzmann Machines¶
Authors: Nair, V., & Hinton, G. E. Year: 2010 ArXiv: https://proceedings.mlr.press/v15/nair10a.html
Summary¶
ReLU introduces the simple but effective max(0, x) activation function. Despite its simplicity, ReLU enabled training of much deeper networks than sigmoid or tanh, becoming the default activation function in modern deep learning.
Key Concepts¶
- Non-saturation for positive values
- Simple computation (max operation)
- Sparsity-inducing properties
- Enables training of very deep networks
- Variants: Leaky ReLU, ELU, GELU
Impact¶
ReLU revolutionized deep learning by solving the vanishing gradient problem. It's used in nearly all modern neural networks and remains the most popular activation function after 15+ years.
Related Papers¶
- Leaky ReLU and other variants
- Understanding the difficulty of training deep feedforward neural networks
Citation:
@inproceedings{nair2010rectified,
title={Rectified linear units improve restricted boltzmann machines},
author={Nair, Vinod and Hinton, Geoffrey E},
booktitle={Proceedings of the 27th international conference on machine learning},
pages={807--814},
year={2010}
}