Skip to content

Random Forests

Authors: Leo Breiman Year: 2001 Citations: 50,000+

Summary

Ensemble of decision trees on random data/feature subsets via bootstrap aggregating. Simple yet powerful algorithm with automatic feature selection and reduced overfitting.

Key Concepts

  • Bagging: Bootstrap samples for diverse training sets
  • Random Feature Selection: Random subset per split
  • Ensemble Aggregation: Majority voting or averaging
  • Out-of-Bag Error: Internal validation
  • Feature Importance: Measures predictive value
  • Parallelization: Independent tree training

Impact

  • 50,000+ citations
  • Dominant algorithm before deep learning
  • Strong baseline on diverse tasks
  • Influenced XGBoost, LightGBM
  • Widely deployed in practice
  • Automatic feature selection
  • Still competitive with deep learning
  • CART (Breiman et al., 1984)
  • Gradient Boosting (Friedman, 2001)
  • XGBoost (Chen & Guestrin, 2016)