Skip to content

Asynchronous Methods for Deep Reinforcement Learning (A3C)

Authors: Volodymyr Mnih, Adrià Puigdomènech Badia, Mehdi Mirza, Alex Graves, et al. Year: 2016 Venue: ICML 2016 Citations: 15,000+

Summary

Introduces asynchronous advantage actor-critic (A3C) algorithm. Multiple agents train asynchronously on different experience, updating shared network parameters. Scales well to parallel computation without experience replay.

Key Concepts

  • Actor-Critic: Separate networks for policy (actor) and value (critic)
  • Asynchronous Training: Parallel agents explore independently
  • Advantage Function: Reduces variance in policy gradient
  • Parallel Computation: Scales to multiple cores/machines
  • Policy Gradient: Direct gradient on policy parameters
  • No Experience Replay: Uses on-policy learning instead

Impact

  • 15,000+ citations
  • Practical algorithm for large-scale RL training
  • Efficient use of parallel computation
  • Inspired distributed RL methods
  • Better scaling than DQN on multi-core systems
  • Foundation for modern distributed RL
  • Important for real-world deployment

Key Results

  • Faster training than DQN through parallelization
  • No need for experience replay (on-policy)
  • Works on diverse control tasks
  • Improved sample efficiency
  • Deep Q-Networks (Mnih et al., 2013)
  • Trust Region Policy Optimization (Schulman et al., 2015)
  • Proximal Policy Optimization (Schulman et al., 2017)
  • Distributed Deep Reinforcement Learning (Espeholt et al., 2018)