Skip to content

Proximal Policy Optimization Algorithms (PPO)

Authors: John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, Oleg Klimov Year: 2017 Venue: arXiv 2017 Citations: 20,000+

Summary

Introduces Proximal Policy Optimization (PPO), simplifying TRPO with clipped objective function. Balances sample efficiency with implementation simplicity. Becomes go-to policy gradient method for deep RL.

Key Concepts

  • Policy Gradient: Direct optimization of policy parameters
  • Clipped Objective: Constrains policy updates without KL divergence
  • Trust Region: Prevents large policy changes
  • Sample Efficiency: Uses few samples per update
  • Stability: Robust to hyperparameter choices
  • Simplicity: Easier to implement than TRPO

Impact

  • 20,000+ citations
  • Became dominant policy gradient algorithm in RL
  • Widely used in robotics and control research
  • Foundation for many modern RL systems
  • Simpler than TRPO while maintaining performance
  • Good baseline for policy gradient methods
  • Used in production RL systems

Key Results

  • State-of-the-art performance on continuous control
  • Better sample efficiency than A3C
  • Easier to tune than TRPO
  • Stable training across diverse tasks
  • Trust Region Policy Optimization (Schulman et al., 2015)
  • Actor-Critic Methods (Konda & Tsitsiklis, 2000)
  • Policy Gradient Theorem (Sutton et al., 2000)
  • Deep Reinforcement Learning from Human Preferences (Christiano et al., 2017)