Skip to content

Momentum Contrast for Unsupervised Visual Representation Learning (MoCo)

Authors: Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, Ross Girshick Year: 2020 Venue: CVPR 2020 Citations: 10,000+

Summary

Introduces Momentum Contrast (MoCo), building large queues of negative samples for contrastive learning. Uses momentum encoder to update queue, enabling large-scale self-supervised visual representation learning.

Key Concepts

  • Contrastive Learning: Maximize agreement between augmented views
  • Momentum Encoder: Slowly updated copy of query encoder
  • Queue of Negatives: Maintains large dictionary of negative samples
  • Consistency: Momentum helps maintain consistency across batches
  • Efficiency: Works with memory-efficient implementation
  • Scalability: Larger negative sample sets improve learning

Impact

  • 10,000+ citations
  • Competing approach to SimCLR for self-supervised learning
  • Better scalability with momentum encoder
  • Influenced many subsequent contrastive methods
  • Foundation for vision pre-training without labels
  • Important for understanding contrastive mechanisms

Key Results

  • Competitive with SimCLR and supervised ImageNet pre-training
  • Better memory efficiency than SimCLR
  • Strong transfer learning performance
  • Generalizes to downstream tasks
  • SimCLR: Contrastive Learning (Chen et al., 2020)
  • Bootstrap Your Own Latent (BYOL, Grill et al., 2020)
  • Contrastive Predictive Coding (van den Oord et al., 2018)
  • Metric Learning Fundamentals (Siamese networks, 2015)