Asynchronous Methods for Deep Reinforcement Learning (A3C)¶
Authors: Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, et al. Year: 2016 Venue: ICML 2016 Citations: 15,000+
Link¶
Summary¶
Introduces asynchronous advantage actor-critic (A3C) algorithm. Multiple agents train asynchronously on different experience, updating shared network parameters. Scales well to parallel computation without experience replay.
Key Concepts¶
- Actor-Critic: Separate networks for policy (actor) and value (critic)
- Asynchronous Training: Parallel agents explore independently
- Advantage Function: Reduces variance in policy gradient
- Parallel Computation: Scales to multiple cores/machines
- Policy Gradient: Direct gradient on policy parameters
- No Experience Replay: Uses on-policy learning instead
Impact¶
- 15,000+ citations
- Practical algorithm for large-scale RL training
- Efficient use of parallel computation
- Inspired distributed RL methods
- Better scaling than DQN on multi-core systems
- Foundation for modern distributed RL
- Important for real-world deployment
Key Results¶
- Faster training than DQN through parallelization
- No need for experience replay (on-policy)
- Works on diverse control tasks
- Improved sample efficiency
Related Papers¶
- Deep Q-Networks (Mnih et al., 2013)
- Trust Region Policy Optimization (Schulman et al., 2015)
- Proximal Policy Optimization (Schulman et al., 2017)
- Distributed Deep Reinforcement Learning (Espeholt et al., 2018)