Skip to content

Gradient Boosting Machines

Authors: Jerome H. Friedman Year: 2001 Citations: 20,000+

Summary

Weak learners sequentially trained to predict residuals via gradient descent on function space. Powerful ensemble technique winning Kaggle competitions.

Key Concepts

  • Residual Learning: Each learner fits previous errors
  • Functional Gradients: Gradients on function space
  • Shrinkage: Learning rate regularization
  • Tree Boosting: Gradient boosting with decision trees
  • Loss Functions: Works with differentiable objectives

Impact

  • 20,000+ citations
  • Foundation for XGBoost, LightGBM, CatBoost
  • Dominant in ML competitions
  • State-of-the-art on tabular data
  • Better than Random Forests on many tasks
  • Widely used in industry
  • AdaBoost (Freund & Schapire, 1997)
  • Bagging (Breiman, 1996)
  • XGBoost (Chen & Guestrin, 2016)