Random Forests¶
Authors: Leo Breiman Year: 2001 Citations: 50,000+
Summary¶
Ensemble of decision trees on random data/feature subsets via bootstrap aggregating. Simple yet powerful algorithm with automatic feature selection and reduced overfitting.
Key Concepts¶
- Bagging: Bootstrap samples for diverse training sets
- Random Feature Selection: Random subset per split
- Ensemble Aggregation: Majority voting or averaging
- Out-of-Bag Error: Internal validation
- Feature Importance: Measures predictive value
- Parallelization: Independent tree training
Impact¶
- 50,000+ citations
- Dominant algorithm before deep learning
- Strong baseline on diverse tasks
- Influenced XGBoost, LightGBM
- Widely deployed in practice
- Automatic feature selection
- Still competitive with deep learning
Related Papers¶
- CART (Breiman et al., 1984)
- Gradient Boosting (Friedman, 2001)
- XGBoost (Chen & Guestrin, 2016)