A Statistical Interpretation of Term Specificity and Its Application in Retrieval¶
Authors: Karen Sparck Jones Year: 1972 Citations: 30,000+
Summary¶
Statistical term weighting for IR. TF-IDF scores based on term frequency and document rarity. Core algorithm for decades, still widely used baseline.
Key Concepts¶
- Term Frequency (TF): How often term appears
- Inverse Document Frequency (IDF): Rarity across corpus
- Vector Space Model: Documents as weighted term vectors
- Cosine Similarity: Relevance measurement
- Probabilistic Foundations: Statistical term importance
Impact¶
- 30,000+ citations
- Foundational IR algorithm for 40+ years
- Influenced BM25 and neural IR
- Still used as baseline and feature
- Basis for full-text search engines
Related Papers¶
- BM25 (Robertson & Zaragoza, 2009)
- Vector Space Model (Salton et al., 1975)
- Dense Retrieval (Karpukhin et al., 2020)