Fine-tuning Fundamentals¶
Overview¶
Fine-tuning continues training a pre-trained model on task-specific data to adapt it. Different from training from scratch (cheaper) or using prompts (limited). Foundation for domain adaptation.
- Approach: Start with pre-trained weights, update on task data
- Cost: 10-100x cheaper than pretraining
- Quality: 5-30% improvement typical
- Trade-off: Requires labeled data vs. few-shot prompting
- When: Have domain-specific data or need specific capabilities
-
The Fine-tuning Spectrum¶
Approaches by Complexity & Cost¶
Complexity & Cost Scale:
- ┌─────────────────────────────────────────────────────────┐
- Cost │
- ▲ │
│ │ │
- Full Fine-tune │
- ██ │
│ │ │
- QLoRA / LoRA │
- ██ │
│ │ │
- In-Context Learning (ICL) │
- ██ │
│ │ │
- Prompt Engineering │
- ██ │
│ │ │
- → Quality
- (cheap) (best quality)
- ┘
Selection by scenario:
Simple task, no data:
- Prompt engineering (free)
Task needs some adaptation:
- In-Context Learning / Few-shot (free)
Task needs better quality:
- LoRA / QLoRA ($10-100)
Task is critical, budget available:
- Full fine-tune ($100-10,000)
-
Fine-tuning Approaches¶
1. Prompt Engineering (Zero-shot & Few-shot)¶
No model changes, just better prompts
Few-shot example:
"""
Examples:
Q: What is the capital of France?
A: Paris
Q: What is the capital of Japan?
A: Tokyo
Q: What is the capital of Brazil?
A: Brasília
"""
Pros:
Free (no compute)
No data needed
Works instantly
No training required
Cons:
Limited to model's knowledge
Hard to get precise behavior
Context window limited
Not reproducible (LLM may vary)
When to use:
- Quick experiments
- No budget
- Task is simple
2. In-Context Learning (ICL) / Few-shot¶
Provide examples in context to teach model
Advanced technique:
"""
You are a sentiment classifier.
Your job is to classify tweets as positive or negative.
Examples:
Tweet: "Love this product! So happy!" → Label: Positive
Tweet: "Terrible experience, never again" → Label: Negative
Tweet: "It's okay, nothing special" → Label: Neutral
Now classify this:
Tweet: "Best purchase ever made!"
Label: """
Model learns from examples without weight updates!
Pros:
No training
Flexible (change examples dynamically)
Free
Fast
Cons:
Limited by context window
Needs good examples
Less effective than fine-tuning
Performance depends on example quality
3. Parameter-Efficient Fine-tuning (LoRA, QLoRA)¶
Update only small adapter weights, not full model
Trade-off:
- Cost: ~$10-100 (small GPU, short time)
- Quality: 90-95% of full fine-tune
- Effort: Low
- Data: Moderate (1000-100K examples)
Example (LoRA):
```python
from peft import LoraConfig, get_peft_model
from transformers import AutoModelForCausalLM
# Load pre-trained model
model = AutoModelForCausalLM.from_pretrained("meta-llama/Llama-2-7b")
# Add LoRA
lora_config = LoraConfig(
r=16,
lora_alpha=32,
lora_dropout=0.05,
bias="none"
)
model = get_peft_model(model, lora_config)
# Fine-tune (only LoRA parameters trained!)
trainer = Trainer(model=model, args=training_args,...)
trainer.train()
# Save only LoRA weights (~100MB instead of 13GB)
model.save_pretrained("lora_model")
When to use:
- Have domain data
- Need better than prompting
- Budget limited
- Quality ~90% acceptable
### 4. Full Fine-tuning
```sql
Update all model weights
Trade-off:
- Cost: ~$100-10,000 (many GPU hours)
- Quality: Best (100%, slightly better than LoRA)
- Effort: High
- Data: Needed (10K-1M examples)
When to use:
- Quality is critical
- Have substantial budget
- Have large labeled dataset
- Can't use parameter-efficient methods
Fine-tuning vs Alternatives¶
Comparison Matrix¶
Approach Cost Time Quality Data Needed Complexity
──────────────────────────────────────────────────────────────────────
Prompting $0 Instant 70% None Low
Few-shot ICL $0 Instant 75% <10 examples Low
LoRA $50 Hours 92% 1K-10K Medium
Full Fine-tune $1K Days 95% 10K-100K High
Continued PT $10K+ Weeks 98% 100K-1M Very High
Retraining $100K+ Months 100% 1M+ Extreme
Cost example (LLaMA 7B):
Fine-tune task: Sentiment classification
- Prompting: $0 (just try)
- Few-shot: $0 (include examples)
- LoRA: ~$20 (1 A100 hour)
- Full fine-tune: ~$500 (10 A100 hours)
- Pre-train: $50,000+ (not worth it!)
Recommendation:
- Start: Prompting (fast, free)
- Iterate: Few-shot (still free)
- If good: LoRA ($20-100)
- If critical: Full fine-tune ($500+)
- Only if very important: Pre-training ($50K+)
-
Types of Fine-tuning¶
1. Instruction Fine-tuning (SFT)¶
Teach model to follow instructions
Data format:
[
{"instruction": "Classify sentiment", "input": "Great product!", "output": "Positive"},
{"instruction": "Summarize", "input": "Long text...", "output": "Summary"},
...
]
Goal:
- Model learns to follow instructions
- Better instruction following
- Generalizes to new instructions
Quality improvement:
- Before: Generic responses
- After: Follows instructions precisely
- Improvement: 10-20%
2. Domain Adaptation¶
Teach model domain-specific knowledge
Example: Medical domain
- Fine-tune on medical texts
- Model learns medical terminology
- Knows disease symptoms, treatments
- Better medical Q&A
- Quality: 15-30% improvement
Approach:
- Collect domain data (medical documents, QA pairs)
- Fine-tune on this data
- Model specializes in domain
- Can then further fine-tune for specific tasks
3. Task-Specific Fine-tuning¶
Optimize for specific task
Examples:
- Sentiment analysis: Fine-tune on sentiment data
- Named entity recognition: Fine-tune on entity tagging data
- Summarization: Fine-tune on summary pairs
- Translation: Fine-tune on parallel texts
Quality:
- General model: 70% accuracy
- Task-specific fine-tune: 90% accuracy
- Improvement: 20%
Cost/benefit:
- Most cost-effective fine-tuning
- High ROI (20% improvement, moderate cost)
-
Data Requirements¶
How Much Data Do You Need?¶
Rule of thumb:
Simple task (classification):
- 100-500 examples minimum
- 5,000-10,000 recommended
- 100,000+ for best quality
Complex task (generation):
- 1,000 examples minimum
- 10,000-50,000 recommended
- 500,000+ for best quality
Example (sentiment analysis):
100 examples:
- LoRA: 75% accuracy (OK)
- Full fine-tune: 80% accuracy
1,000 examples:
- LoRA: 88% accuracy (good)
- Full fine-tune: 91% accuracy
10,000 examples:
- LoRA: 93% accuracy (excellent)
- Full fine-tune: 95% accuracy
100,000 examples:
- LoRA: 96% accuracy (near-optimal)
- Full fine-tune: 97% accuracy
Diminishing returns:
- First 1K examples help most
- Each 10x increase in data: ~3-5% improvement
- After 100K examples: marginal gains
Data Quality > Quantity¶
Important insight:
1,000 high-quality examples > 100,000 low-quality examples
Quality characteristics:
- Accurate labels (no noise)
- Representative of your task
- Well-formatted
- No duplicates/near-duplicates
- Consistent annotation
Example:
10K noisy examples (50% mislabeled):
- Model learns noise
- Accuracy: 70%
1K clean examples:
- Model learns signal
- Accuracy: 85%
Lesson:
- Spend time on data quality!
- Small clean dataset beats large noisy dataset
-
Fine-tuning Process¶
Basic Pipeline¶
Step 1: Data Preparation
- Collect domain/task data
- Format in expected structure
- Split: Train/Val/Test (80/10/10)
- Quality check
- Time: Hours to days
Step 2: Model Selection
- Choose base model
- Consider size (7B vs 70B)
- Consider pre-training quality
- Time: Minutes
Step 3: Fine-tuning
- Choose method (LoRA vs Full)
- Set hyperparameters
- Train on data
- Time: Hours to days
Step 4: Evaluation
- Validate on held-out data
- Compare to baseline
- Check for overfitting
- Time: Hours
Step 5: Iteration
- Analyze errors
- Collect more data
- Adjust hyperparameters
- Retrain
- Time: Days to weeks
Step 6: Deployment
- Merge weights (if LoRA)
- Package model
- Deploy to production
- Time: Hours
Common Mistakes¶
Mistake 1: Insufficient Data¶
Bad: Fine-tune on 100 examples
- Model overfits
- Poor generalization
- Not worth the effort
Good: Collect 1,000+ examples
- Model learns patterns
- Generalizes better
- Worth the investment
Mistake 2: Not Evaluating on Held-out Data¶
Bad: Train on all data, evaluate on training data
- Appears 95% accurate
- Actually 70% on new data (overfitting!)
- False confidence
Good: Split data, evaluate on test set
- Honest evaluation
- Can detect overfitting
- Real performance estimate
Mistake 3: Training Too Long¶
Bad: Train for 100 epochs
- Overfits to training data
- Validation accuracy decreases
- Model memorizes instead of learns
Good: Use early stopping
- Monitor validation loss
- Stop when it starts increasing
- Get best model without overfitting
Implementation:
```python
from transformers import EarlyStoppingCallback
trainer = Trainer(
model=model,
callbacks=[EarlyStoppingCallback(early_stopping_patience=3)]
)
trainer.train()
```
Key Takeaways¶
Fine-tuning vs Prompting: 10-30% quality improvement LoRA: Best cost-effectiveness (90% quality, low cost) Data quality > quantity: 1K clean beats 100K noisy Start simple: Prompting → Few-shot → LoRA → Full Evaluate properly: Always use held-out test set
-
Related Notes in Finetuning Subdirectory¶
- [Instruction Tuning & Sft](/01-modeling/02-training/finetuning/(instruction-tuning-sft/) - Teaching instructions
- Dpo (Direct Preference Optimization) - Learning from preferences
- Rlhf - Reinforcement learning approach (in main Modeling)
- Domain Specific Fine Tuning - Specialized adaptation
- Fine Tuning Best Practices - Tips and tricks