Skip to content

Mastering the Game of Go with Deep Neural Networks and Tree Search (AlphaGo)

Authors: David Silver, Aja Huang, Chris J. Maddison, Arthur Guez, Laurent Sifre, et al. Year: 2016 Venue: Nature 2016 Citations: 20,000+

Summary

Introduces AlphaGo, combining deep neural networks with Monte Carlo Tree Search (MCTS) for Go. Uses policy network for action selection and value network for board evaluation. First AI system to defeat professional Go players.

Key Concepts

  • Policy Network: Predicts probability of moves via deep CNN
  • Value Network: Estimates board evaluation/winning probability
  • Monte Carlo Tree Search: Explores game tree with network guidance
  • Self-Play: Trains via playing against itself
  • Combination of Learning and Search: Integrates learned policy/value with planning
  • Go Complexity: Billion-scale action space requiring smart exploration

Impact

  • 20,000+ citations
  • Historic milestone: AI defeats top Go players
  • Demonstrated power of combining learning and search
  • Inspired AlphaGo Zero (learning without human data)
  • Showed scalability of deep RL to complex domains
  • Influenced game-playing AI research
  • Demonstrated AI capabilities to broad audiences

Key Results

  • Defeated Lee Sedol (top Go player) 4-1
  • Superhuman performance on complex game
  • Combination of MCTS and deep learning effective
  • Generalized to new positions despite training diversity
  • Deep Q-Networks (Mnih et al., 2013)
  • AlphaGo Zero (Silver et al., 2017)
  • AlphaZero (Silver et al., 2018)
  • Monte Carlo Tree Search (Kocsis & Szepesvári, 2006)