Enriching Word Vectors with Subword Information (FastText)¶
Authors: Piotr Bojanowski, Edouard Grave, Armand Joulin, Tomas Mikolov Year: 2017 (arXiv 2016) Venue: TACL 2017 Citations: 15,000+
Link¶
Summary¶
Introduces FastText, representing words as sums of character n-grams instead of single word vectors. Improves handling of morphologically rich languages, out-of-vocabulary words, and achieves competitive performance with faster training.
Key Concepts¶
- Subword Information: Uses character n-grams instead of whole words
- Morphological Awareness: Captures morphological structure
- Out-of-Vocabulary Handling: Can represent any word from n-grams
- Fast Training: Efficient computation via hash bucketing
- Language Diversity: Better for morphologically rich languages
- Compositional Semantics: Words built from semantic subword units
Impact¶
- 15,000+ citations
- Improved word embeddings for languages with complex morphology
- Better handling of rare and out-of-vocabulary words
- Faster training than Word2Vec
- Foundation for modern subword-based models (BPE, WordPiece)
- Influenced BERT and Transformers tokenization
- Practical and widely used for embeddings
Key Results¶
- Better performance on similarity tasks
- Good handling of morphologically complex languages
- Faster training than Word2Vec on some settings
- Competitive with other word embedding methods
Related Papers¶
- Word2Vec (Mikolov et al., 2013)
- GloVe (Pennington et al., 2014)
- Byte Pair Encoding (Sennrich et al., 2016)
- BERT (Devlin et al., 2018)