Chunking & Document Preparation¶
Overview¶
Chunking splits large documents into smaller pieces for embedding and retrieval. Document preparation cleans and structures documents for optimal retrieval. These are critical for RAG quality - garbage in = garbage out!
- Chunking: Split documents into ~500-2000 token pieces
- Strategy: Size, overlap, semantic awareness
- Impact: Directly affects retrieval quality (critical!)
- Trade-off: Smaller chunks = more granular, larger chunks = more context
The Problem¶
Why Chunking Matters¶
Scenario: Legal document 50KB (50,000 tokens)
Option 1: Store entire document as single embedding
- Embedding: [single vector for whole doc]
- Problem: Vector averages out all semantics
- Query about "section 3.2" retrieves entire 50KB doc
- LLM gets overwhelmed with irrelevant content
- Quality suffers
- Result: Poor retrieval quality
Option 2: Split into chunks of 500 tokens each
- Chunks: 100 vectors (one per chunk)
- Problem solved: Query about "section 3.2" retrieves only that section!
- LLM gets focused context
- Quality improves
- Result: Much better retrieval quality!
Example impact:
Without chunking: 60% accuracy (retrieves too much noise)
With chunking: 90% accuracy (retrieves focused context)
Impact: 30% accuracy improvement from better chunking!
-
Chunking Strategies¶
1. Fixed-Size Chunking¶
Split documents into fixed-size pieces
Configuration:
- Chunk size: 500 tokens
- Overlap: 50 tokens (prevent losing context)
- Strategy: Simple split
Example:
Document:
[Token 1-500] [Token 451-950] [Token 901-1400] [Token 1351-1850]...
- Chunk 1 └─ Chunk 2 └─ Chunk 3 └─ Chunk 4
Chunk 1: Tokens 1-500
Chunk 2: Tokens 451-950 (50 tokens overlap)
Chunk 3: Tokens 901-1400 (50 tokens overlap)
...
Pros:
Simple to implement
Fast
Predictable
Cons:
Splits can break sentences/meaning
No semantic awareness
Overlap wastes tokens
2. Semantic Chunking¶
Split at sentence/paragraph boundaries to preserve meaning
Algorithm:
- Calculate embedding for each sentence
- When embedding distance > threshold, split
- Result: Chunks semantically cohesive
Example:
Sentence embeddings:
S1: "The cat sat on the mat." [0.12, -0.45,...]
S2: "It was a sunny day." [0.13, -0.44,...] (similar to S1: 0.98)
S3: "Photosynthesis is the process..." [0.78, 0.34,...] (different: 0.15) ← SPLIT HERE
S4: "Plants convert light to energy." [0.79, 0.33,...] (similar to S3: 0.99)
Chunks:
Chunk 1: "The cat sat on the mat. It was a sunny day."
Chunk 2: "Photosynthesis is the process... Plants convert light to energy."
Pros:
Respects sentence boundaries
Semantically coherent chunks
Better retrieval quality
Cons:
More complex implementation
Slower (computes embeddings per sentence)
3. Recursive Chunking¶
Start with large chunks, recursively split if too big
Algorithm:
- Try to chunk on sentence boundaries
- If sentence too long, chunk on paragraph boundaries
- If still too long, chunk on line breaks
- If still too long, split by character
- Result: Balanced chunks with good boundaries
Use case: Heterogeneous documents (PDFs, code, tables)
Example:
Document
- Try sentence split: Average 200 tokens (OK)
- Try paragraph split: Average 800 tokens (too big!)
- Try sentence split: Average 200 tokens (use this)
- Result: Sentence-level chunks
Code document:
- Try line split: Average 50 tokens (too small, wasteful)
- Try function split: Average 300 tokens (good!)
- Result: Function-level chunks
Pros:
Flexible for mixed content
Respects document structure
Near-optimal chunk sizes
Cons:
More complex
Slower for large documents
Chunk Size Selection¶
Optimal Size¶
Chunk size trade-off:
Small chunks (100-300 tokens):
More specific retrieval
Faster embedding
Less context
More chunks (slower search)
Medium chunks (500-800 tokens):
Good balance
Enough context for LLM
Not too many vectors to search
- Recommended!
Large chunks (1000-2000 tokens):
Maximum context
Fewer chunks
Less granular retrieval
More noise in results
Empirical results (on QA tasks):
Chunk size Retrieval Acc. Generation Quality
──────────────────────────────────────────────
256 tokens 0.72 0.65
512 tokens 0.85 0.82
1024 tokens 0.82 0.90
2048 tokens 0.78 0.92
Observation:
- 512 tokens: Best retrieval
- 1024 tokens: Good retrieval + good generation
- 2048 tokens: Poor retrieval but good generation
- Recommendation: 512-1024 tokens
Why 512-1024 is optimal:
- Retrieves specific information (good)
- Provides enough context for LLM (good)
- Manageable vector search (good)
- Not too much noise (good)
Overlap¶
Overlap: Repeat tokens between chunks
Purpose: Don't lose context at boundaries
Example (chunk size 500, overlap 50):
Document: [1-500] [451-950] [901-1400]
Tokens: 1-----500
451-------950
901--------1400
Overlap benefits:
- Preserve context across chunks
- Relevant information not at boundary
- Query might match partial context
- Improved retrieval quality
Overlap cost:
- Duplicate vectors stored (more memory)
- Slower search (more vectors)
- ~10% overhead per overlap
Typical: Overlap = 10% of chunk size
- Chunk 500: Overlap 50 tokens (10%)
- Chunk 1000: Overlap 100 tokens (10%)
- Balance: Quality vs efficiency
-
Document Preprocessing¶
Cleaning¶
Remove noise before chunking:
1. Remove HTML/XML tags
Before: "<p>The <b>cat</b> sat...</p>"
After: "The cat sat..."
2. Remove extra whitespace
Before: "The cat\n\n\nsat"
After: "The cat sat"
3. Remove headers/footers
Before: "Page 1| Document Title\n\nThe cat sat..."
After: "The cat sat..."
4. Normalize encoding
Before: "café" (wrong encoding)
After: "café" (correct UTF-8)
5. Remove duplicates
- Same content in multiple places → keep once
Quality impact:
- Cleaning alone: ~5-10% improvement
Metadata Extraction¶
Extract and preserve important information:
Example: Legal document
Extract:
- Document type: "Contract"
- Date: "2024-01-15"
- Parties: ["Company A", "Company B"]
- Section: "3.2 Payment Terms"
- Page number: 5
Store with chunk:
{
"id": "doc_chunk_42",
"text": "Payment is due within 30 days...",
"embedding": [0.12, -0.45,...],
"metadata": {
"type": "Contract",
"date": "2024-01-15",
"section": "3.2",
"page": 5,
"parties": ["Company A", "Company B"]
}
}
Use in retrieval:
- Filter by document type
- Filter by date range
- Filter by section
- More precise retrieval!
Quality impact:
- Metadata filtering: ~10-20% improvement
-
Special Cases¶
Code Documents¶
Problem: Code has structure (functions, classes)
Poor chunking:
- Split code mid-function
- Imports separated from usage
- Related code in different chunks
Better chunking:
- Split at function/class boundaries
- Keep imports with code
- Preserve semantic structure
Example:
def calculate_tax(income):
"""Calculate tax based on income."""
if income < 50000:
return income * 0.1
else:
return income * 0.2
Chunk this unit intact!
Don't split function in half.
Implementation:
```python
def chunk_code(code_str, language="python"):
"""Chunk code by function/class boundaries"""
if language == "python":
# Parse AST
tree = ast.parse(code_str)
chunks = []
for node in tree.body:
if isinstance(node, (ast.FunctionDef, ast.ClassDef)):
chunk = get_node_source(node)
chunks.append(chunk)
return chunks
Quality impact:
- Code-aware chunking: 15-30% improvement for code QA
### Tables & Structured Data
Problem: Tables don't chunk well
Table:
| Name | Age | City |
|---|---|---|
| Alice | 30 | NYC |
| Bob | 25 | LA |
Naive chunk: Splits table rows → Lost structure!
Solution: Keep tables intact + add context
Chunk:
Sales data for Q3 2024:
| Product| Revenue| Growth|
|---------|---------|--------|
| Widget| $1M| +20%|
| Gadget| $2M| +15%|
Implementation:
def chunk_with_tables(html_content):
"""Chunk while preserving tables"""
tables = extract_tables(html_content)
text = remove_tables(html_content)
chunks = []
for chunk in chunk_text(text):
# Add nearby tables to chunk
related_tables = find_nearby_tables(chunk, tables)
chunk_with_tables = chunk + related_tables
chunks.append(chunk_with_tables)
return chunks
Quality impact:
- Table-aware chunking: 20-40% improvement for table QA
```
Key Takeaways¶
Chunking: Critical for RAG quality (garbage in = garbage out) Optimal size: 512-1024 tokens (balances specificity + context) Overlap: ~10% prevents losing context at boundaries Preprocessing: Cleaning + metadata extraction matters Domain-aware: Code, tables, PDFs need special handling
-
Related Notes in RAG Subdirectory¶
- Rag Fundamentals - Context for chunking importance
- Retrieval Strategies - How chunking affects retrieval
- [Vector Databases & Embeddings](/01-modeling/03-inference/03-knowledge-integration/rag/(vector-databases-embeddings/) - Where chunks are stored