GloVe: Global Vectors for Word Representation¶
Authors: Jeffrey Pennington, Richard Socher, Christopher D. Manning Year: 2014 Venue: EMNLP 2014 Citations: 25,000+
Link¶
Summary¶
Global log-bilinear regression model combining global matrix factorization with local context. Uses co-occurrence statistics from corpus to learn embeddings capturing both global structure and local context.
Key Concepts¶
- Co-occurrence Matrix: Global word-word statistics from corpus
- Weighted Least Squares: Balances rare and frequent pairs
- Vector Composition: Captures semantic and syntactic relationships
- Efficiency: Unsupervised training on co-occurrence statistics
Impact¶
- 25,000+ citations
- Superior semantic performance vs Word2Vec on many tasks
- Demonstrates importance of global corpus statistics
- Widely adopted in NLP research
- Still used as baseline embeddings
Related Papers¶
- Word2Vec (Mikolov et al., 2013)
- Distributed Representations (Mikolov et al., 2013)
- FastText (Bojanowski et al., 2017)