Skip to content

GloVe: Global Vectors for Word Representation

Authors: Jeffrey Pennington, Richard Socher, Christopher D. Manning Year: 2014 Venue: EMNLP 2014 Citations: 25,000+

Summary

Global log-bilinear regression model combining global matrix factorization with local context. Uses co-occurrence statistics from corpus to learn embeddings capturing both global structure and local context.

Key Concepts

  • Co-occurrence Matrix: Global word-word statistics from corpus
  • Weighted Least Squares: Balances rare and frequent pairs
  • Vector Composition: Captures semantic and syntactic relationships
  • Efficiency: Unsupervised training on co-occurrence statistics

Impact

  • 25,000+ citations
  • Superior semantic performance vs Word2Vec on many tasks
  • Demonstrates importance of global corpus statistics
  • Widely adopted in NLP research
  • Still used as baseline embeddings
  • Word2Vec (Mikolov et al., 2013)
  • Distributed Representations (Mikolov et al., 2013)
  • FastText (Bojanowski et al., 2017)