Token Efficiency & Compression¶
Overview¶
Token Efficiency measures how many tokens are needed to represent text. Lower is better (cheaper inference, faster training). Compression techniques reduce token count without sacrificing quality.
- Goal: Minimize tokens per word
- Current state: 1.1-1.3 tokens/word typical
- Improvement potential: 5-15% reduction possible
- Impact: Direct cost savings (billing per token)
-
Measuring Token Efficiency¶
Token Count Analysis¶
Text: "The quick brown fox jumps over the lazy dog"
Method 1: Count directly
- Words: 9
- Tokens (GPT-3): 11
- Efficiency: 11/9 = 1.22 tokens/word
Method 2: Large corpus
- 1M word document
- Tokens: 1.15M
- Efficiency: 1.15 tokens/word
Implementation:
import tiktoken
def analyze_efficiency(text, model="gpt-3.5-turbo"):
enc = tiktoken.encoding_for_model(model)
tokens = enc.encode(text)
words = text.split()
efficiency = len(tokens) / len(words)
print(f"Text: {text}")
print(f"Words: {len(words)}")
print(f"Tokens: {len(tokens)}")
print(f"Efficiency: {efficiency:.2f} tokens/word")
return efficiency
# Test various texts
texts = [
"Hello world",
"The quick brown fox jumps over the lazy dog",
"Pneumonoultramicroscopicsilicovolcanoconiosis",
"2023-01-15 10:30:45",
"int main() { printf(\"hello\"); }",
]
for text in texts:
analyze_efficiency(text)
Efficiency by Text Type¶
Text Type Tokens/Word Reason
──────────────────────────────────────────────────
Common English 1.10 Frequent words
Technical text 1.20 Jargon, abbreviations
Code 1.30 Symbols, identifiers
Dates/Numbers 1.50 Each digit can be token
Emoji 2.00+ Multiple tokens per emoji
Mixed language 1.50-2.00 Language switching
Asian languages 3.00+ No spaces, character-based
Example breakdown:
"Python 3.11 released on 2023-10-02"
Tokens: ["Python", "▁3", ".", "11", "▁released", "▁on", "▁2023", "-", "10", "-", "02"]
Count: 11 tokens for 6 words = 1.83 tokens/word
Why: Numbers expand (3,., 1, 1 are separate)
-
Techniques to Improve Efficiency¶
1. Better Tokenizer Training¶
Observation: Tokenizer quality affects efficiency
Poor tokenizer (untrained):
- "wonderful" → ["wo", "nde", "rfu", "l"] (4 tokens)
- Arbitrary splits
- Efficiency: Bad
Good tokenizer (well-trained):
- "wonderful" → ["wonderful"] (1 token)
- Learned this is common word
- Efficiency: Good
How to improve:
- Train on representative data (your domain)
- Larger vocab size (32K → 64K)
- More training data
- Optimize for your use case
Cost/benefit:
- Better tokenizer: 5-10% fewer tokens
- Training cost: 1-2 hours
- Training cost per token saved: Easily worth it!
2. Domain-Specific Tokenization¶
Observation: Different domains have different efficiency
Generic tokenizer on medical text:
- "COVID-19" → ["COVID", "-", "19"] (3 tokens)
- "mRNA" → ["m", "RNA"] (2 tokens)
- Inefficient!
Medical domain tokenizer:
- "COVID-19" → ["COVID-19"] (1 token)
- "mRNA" → ["mRNA"] (1 token)
- Efficient!
Efficiency gain:
- Generic: 1.35 tokens/word (medical text)
- Domain-specific: 1.15 tokens/word
- Savings: 15% fewer tokens!
Implementation:
# Train tokenizer on medical corpus
spm.SentencePieceTrainer.train(
input='medical_texts.txt', # Domain data!
model_prefix='medical_sp',
vocab_size=32000,
)
3. Aggressive Merging¶
More merges in BPE = fewer tokens
BPE training:
- Standard: 50K vocab (typical)
- Aggressive: 128K vocab (larger)
- Larger vocab = fewer merged tokens!
Trade-off:
32K vocab:
- Smaller tokenizer
- Fewer tokens per document
- But: Some word splitting
64K vocab:
- Larger tokenizer file
- Even fewer tokens per document
- More memory for embedding layer
- Trade-off: Size vs efficiency
128K vocab:
- Very large tokenizer
- Near-complete word coverage
- Minimal splitting
- Embedding layer becomes expensive
Recommendation:
- Single language: 32-50K (sweet spot)
- Multilingual: 64K (worth the cost)
- Extreme efficiency: 128K (only if critical)
4. Special Token Handling¶
Observation: Special characters, numbers expand tokens
Bad:
- "2023-01-15" → ["2023", "-", "01", "-", "15"] (5 tokens!)
- "hello@world.com" → ["hello", "@", "world", ".", "com"] (5 tokens)
- Inefficient!
Better: Pre-process to collapse special patterns
import re
def preprocess_for_efficiency(text):
"""Collapse common patterns"""
# Dates: 2023-01-15 → <date>
text = re.sub(r'\d{4}-\d{2}-\d{2}', '<date>', text)
# Emails: test@example.com → <email>
text = re.sub(r'\S+@\S+', '<email>', text)
# URLs: https://example.com → <url>
text = re.sub(r'https?://\S+', '<url>', text)
# Numbers: 12345 → <num>
text = re.sub(r'\b\d{4,}\b', '<num>', text)
return text
# Test
text = "Email me at test@example.com on 2023-01-15. Visit https://example.com"
processed = preprocess_for_efficiency(text)
print(f"Original tokens: {len(tokenizer.encode(text))}")
print(f"Processed tokens: {len(tokenizer.encode(processed))}")
# Likely savings
Compression Techniques¶
Prompt Compression¶
Observation: Long prompts waste tokens
Standard prompt (50+ tokens):
You are a helpful assistant. Answer questions accurately.
Today is 2024-08-08.
The user asks:...
Compressed prompt (30 tokens, 40% savings!):
Answer accurately.
Today: 2024-08-08
User:...
Technique: Remove unnecessary words
- Remove: "You are a", "The user"
- Keep: Essential information
- Result: Shorter, same meaning
Implementation:
def compress_prompt(prompt):
"""Remove redundant words"""
compressions = {
"You are a helpful assistant": "Be helpful",
"Answer the question": "Answer",
"I will ask you": "",
"The user asks": "User asks",
}
for original, replacement in compressions.items():
prompt = prompt.replace(original, replacement)
return prompt.strip()
Context Compression¶
Technique: Summarize long context before passing to model
Problem:
- User provides 10K token context
- Model processes all 10K tokens
- Expensive and slow!
Solution: Compress context first
def compress_context(long_context, target_tokens=1000):
"""Compress context to target size"""
# Option 1: Extractive (keep most important sentences)
sentences = long_context.split('.')
# Score sentences by keyword importance
scores = []
for sentence in sentences:
# Simple: Count words (better: use TF-IDF)
score = len(sentence.split())
scores.append((sentence, score))
# Keep top-scoring sentences until target tokens
selected = []
tokens_so_far = 0
for sentence, score in sorted(scores, key=lambda x: x[1], reverse=True):
tokens = len(sentence.split())
if tokens_so_far + tokens <= target_tokens:
selected.append(sentence)
tokens_so_far += tokens
return '.'.join(selected)
# Option 2
# Use smaller model to summarize
Result:
- Original: 10K tokens
- Compressed: 1K tokens
- Savings: 90%!
- Trade-off: Some context loss, but often acceptable
-
Efficiency Improvements by Technique¶
Technique Improvement Effort Cost/Benefit
─────────────────────────────────────────────────────────────────
Better tokenizer 5-10% Medium High
Domain-specific vocab 10-15% High High
Special token handling 3-5% Low Medium
Larger vocab size 2-5% Low Medium
Prompt compression 10-20% Low High
Context compression 20-90% Medium High (depends on use)
Combined (all) 30-50% High Very high
Real-world example (QA system):
Without optimization:
- Question: 50 tokens
- Context: 2000 tokens
- Total: 2050 tokens per query
With optimization:
- Question (compressed): 40 tokens
- Context (compressed): 800 tokens
- Total: 840 tokens per query
- Savings: 59%!
Cost impact (at $0.003 per 1K tokens):
- 1M queries before: 1M × 2050 tokens × $0.000003 = $6,150
- 1M queries after: 1M × 840 tokens × $0.000003 = $2,520
- Savings: $3,630 (59%)
Best Practices for Efficiency¶
Do's¶
Train tokenizer on representative data
Use domain-specific vocabulary
Compress prompts when possible
Monitor token efficiency
Test efficiency before deployment
Compress context when feasible
Don'ts¶
Don't use generic tokenizer for specialized domain
Don't include unnecessary text in prompts
Don't ignore token compression opportunities
Don't use too-small vocab (inefficient splitting)
Don't use too-large vocab (embedding layer overhead)
Key Takeaways¶
Measure efficiency: Tokens per word is key metric Domain tokenizer: 10-15% savings for specialized domains Compression: Summarize long context for 20-90% savings Special tokens: Collapse patterns for 3-5% savings ROI: Token savings directly reduce costs
-
Related Notes in Tokenization Subdirectory¶
- 00 Tokenization Fundamentals - Overview
- 01 Bpe (Byte Pair Encoding) - Standard tokenization method
- 02 Sentencepiece - Modern alternative
- 05 Multilingual Tokenization - Language-specific efficiency
- 04 Tokenization Best Practices - Tips and pitfalls