Skip to content

Scaling Laws & Optimal Allocation

Overview

Scaling Laws are empirical relationships showing how model/dataset/compute scale affects performance. Optimal Allocation determines how to split a compute budget between parameters, data, and training time for best performance.

  • Foundational Works: "Scaling Laws for Neural Language Models" (Kaplan et al., 2020), "Chinchilla" (Hoffmann et al., 2022)
  • Key Finding: Optimal compute allocation: D ≈ 20N (20x more data tokens than parameters)
  • Previous Assumption: Parameters >> Data (turned out wrong!)
  • Modern Understanding: Balanced scaling of model and data is optimal
  • Practical Impact: Design 7B model? Plan for 140B tokens of training data!

-

The Scaling Law Relationships

Basic Scaling Laws (Kaplan et al. 2020)

Loss as function of model size (N):
L(N) = a × N^α + other terms
where α ≈ -0.06 to -0.08 (depends on metric)

Interpretation:
- Double model size → ~10% better performance
- 10x model size → ~30-40% better performance
- Diminishing returns, but consistent pattern

Empirical formula:
L ≈ L₀ + A/N^α + B/D^β

Where:
 - N = Number of parameters
 - D = Dataset size (tokens)
 - A, B, L₀ = Constants (fitted from data)
 - α ≈ 0.07 (parameter scaling exponent)
 - β ≈ 0.16 (data scaling exponent)

Key observation:
 - β > α: Data helps more than parameters!
 Doubling data is better than doubling model size!

Chinchilla Scaling (Hoffmann et al. 2022)

Assumption (Kaplan): Use compute to add parameters
 - GPT-3: 175B parameters
 - Trained on 300B tokens
 - Ratio: 1B tokens per parameter (N >> D)

Problem: Far from optimal!

New finding (Chinchilla):
Compute-optimal allocation: D ≈ 20N
 - Train each parameter on ~20 tokens
 - Not 2 tokens (Kaplan), not 1 token (Chinchilla limit)
 - Balanced scaling

Implication:
 - GPT-3 (175B params, 300B tokens): Undertrained!
 - Optimal GPT-3 equivalent: 175B params, 3.5T tokens
 - Or: 70B params, 1.4T tokens (better compute efficiency)

Real data:
 - GPT-3: Over-parameterized, under-trained
 - Chinchilla (70B): Same compute, better quality
 - LLaMA 2 70B: Even better (more tokens, better optimized)
 - Modern models follow D ≈ 20N rule!

-

Scaling Laws in Practice

Parametric Scaling Relationship

Loss(N, D) = E/N^α + F/D^β + L_residual

Coefficients from LLaMA training:
 - E ≈ 10 (parameter scaling)
 - F ≈ 160 (data scaling)
 - α ≈ 0.065
 - β ≈ 0.169
 - L_residual ≈ 1.69 (irreducible loss)

Practical formula:
Loss ≈ 10/N^0.065 + 160/D^0.169 + 1.69

Example: How does model size affect loss?

Model size Loss (at D=1T)
──────────────────────────────
1B parameters 10/1 + 160/1000^0.169 + 1.69 ≈ 13.3
7B parameters 10/7 + 160/1000^0.169 + 1.69 ≈ 3.8
70B parameters 10/70 + 160/1000^0.169 + 1.69 ≈ 2.1

Key: Loss improves predictably with model size!
 Can forecast: "If we scale to 100B, expect loss ≈ X"

Compute Budget Allocation

Fixed compute budget C (in FLOPs):

C ≈ 6ND (3x for forward, 2x for gradient computation, etc.)

Question: How to allocate C between N (parameters) and D (data)?

Options:
1. Big model, less data: 100B params, 100B tokens
2. Balanced: 70B params, 1400B tokens
3. Small model, more data: 7B params, 14000B tokens

Which is best?

Loss comparison:

Option 1 (100B params, 100B tokens):
 - Loss ≈ 10/100 + 160/100 + 1.69 ≈ 3.28

Option 2 (70B params, 1400B tokens):
 - Loss ≈ 10/70 + 160/1400 + 1.69 ≈ 2.44

Option 3 (7B params, 14000B tokens):
 - Loss ≈ 10/7 + 160/14000 + 1.69 ≈ 3.2

Winner: Option 2 (Chinchilla scaling)!
Lesson: Balanced allocation better than extremes
 70B with 1.4T tokens better than 100B with 100B tokens

-

Optimal Compute Allocation

The Chinchilla Hypothesis

Chinchilla paper findings:

Compute-optimal frontier:
 - Given compute budget C
 - Optimal: D ≈ 20N (train on 20 tokens per parameter)
 - This maximizes model quality

Allocation formula:
 - N_optimal = (C / 6 / 20)^(1/2) × some_constant
 - D_optimal ≈ 20 × N_optimal
 - Result: Balanced parameter and data allocation

Real example:
 - Compute budget: 10^21 FLOPs
 - Chinchilla allocation: 70B params, 1.4T tokens
 - Kaplan (old): 100B params, 300B tokens
 - Chinchilla is compute-optimal!

LLaMA validation:
 - LLaMA 70B: Trained on 1.4T tokens (Chinchilla aligned)
 - Quality: Better than GPT-3 (175B, 300B tokens)
 - Confirmation: Chinchilla scaling is correct!

Compute Allocation Curves

For a fixed compute budget:

Trade-off: Parameters vs Data Tokens

 Parameters
 ↑
 Large| (100B, 100B)
 model| X
| 70B × 1.4T
 Small| 
 model| (7B, 14B)
 - → Data Tokens

The "Chinchilla curve":
 - Shows compute-optimal points
 - Doubling compute: 70B→100B params, 1.4T→2T tokens
 - Near-diagonal line in log-log space
 - Key: D/N ratio ≈ 20 is stable!

-

Practical Application: Designing Your Model

Step 1: Compute Budget

Question: How much compute do you have?

Typical budgets:
 - Startup: 10^18 FLOPs (~1 day on 8×A100)
 - Research lab: 10^20 FLOPs (~100 days on 1000×A100)
 - Google-scale: 10^23 FLOPs+ (months on massive clusters)

Assume: Research lab with 10^20 FLOPs available

Step 2: Calculate Optimal Model Size

C ≈ 6ND (roughly)

For optimal Chinchilla scaling:
 - D ≈ 20N
 - C ≈ 6N(20N) = 120N²
 - N_optimal = √(C/120)

Example (C = 10^20):
 - N = √(10^20 / 120) ≈ 10^9.5 ≈ 30B parameters
 - D = 20 × 30B ≈ 600B tokens
 - Result: Build 30B model, train on 600B tokens

Step 3: Hardware & Time Estimate

FLOPs vs Real Time:

Model size: 30B
Training data: 600B tokens
Sequence length: 4K

Total FLOPs: 6 × 30B × 600B ≈ 10^20

Hardware: 100 × A100 GPUs
 - A100 performance: 312 TFLOPS (float32)
 - Total: 31 PFLOPS
 - Efficiency: ~70% = 22 PFLOPS realistic

Wall-clock time:
 - 10^20 FLOPs / (22 × 10^15 FLOPS/sec)
 - ≈ 4.5 × 10^3 seconds
 - ≈ 75 minutes
 - So: 100 GPUs, ~1.25 hours to train 30B model on 600B tokens!

With 10 GPUs:
 - Time: 75 minutes × 10 ≈ 750 minutes
 - ≈ 12.5 hours
 - Reasonable for research!

Empirical Validation

Comparing Models by Efficiency

Model Params Tokens Compute (FLOPs) Loss
────────────────────────────────────────────────────
GPT-3 175B 300B 3×10^23 3.0
Chinchilla 70B 1.4T 1.2×10^23 2.6
LLaMA 70B 70B 1.4T 1.2×10^23 2.55
Chinchilla-12B 12B 240B 2×10^20 3.8

Observation:
 - GPT-3: Higher loss despite 2.5x more compute
 - Chinchilla: Better quality with less compute
 - LLaMA: Slightly better than Chinchilla
 - Validation: D ≈ 20N allocation is optimal

-

Advanced Topics

LLaMA's Refinement

LLaMA improvements over Chinchilla:

1. More compute allocated to training
 - 7B: 1T tokens (same as Chinchilla)
 - 13B: 1T tokens (Chinchilla would use ~260B)
 - 70B: 1.4T tokens (Chinchilla would use ~1.4T)
 - Effect: Higher quality for same tokens

2. Optimized tokenization
 - Better vocabulary coverage
 - Fewer tokens needed for same information

3. Better training procedure
 - Improved optimizer
 - Better learning rate schedule
 - Effect: Faster convergence

Result: LLaMA slightly outperforms Chinchilla predictions
 But scaling law exponents (α, β) similar!

IsoFLOP Curves

IsoFLOP: "Same compute, different allocation"

For fixed compute (10^20 FLOPs):

Model Size Tokens Quality
──────────────────────────────
1B 10T Poor (~4.5)
7B 1.4T Good (~2.8)
70B 140B Good (~2.4)
200B 50B Worse (~2.5)

Curve shape (in log space):
 - U-shaped: extreme allocations bad, middle optimal
 - Minimum at ≈70B/1.4T (Chinchilla point)

Insight:
 - Too-small models waste data potential
 - Too-large models waste parameter potential
 - Sweet spot: Chinchilla scaling

-

Practical Guidelines

For Model Selection

You have 1 week of training time on 8×A100:

Step 1: Calculate FLOPs budget
 - 8 A100s × 312 TFLOPS × 86400 sec/day × 7 days × 0.7 efficiency
 - ≈ 1.2 × 10^20 FLOPs available

Step 2: Apply Chinchilla
 - N ≈ √(1.2 × 10^20 / 120) ≈ 30B parameters
 - D ≈ 20 × 30B = 600B tokens

Step 3: Train
 - Build or download 30B model
 - Prepare 600B tokens of data
 - Train for 7 days
 - Result: Optimal quality for your budget

For Organization Scale

Scale Budget Model Size Tokens
──────────────────────────────────────────────────
Startup 10^18 FLOPs 1-3B 20-60B
Research Lab 10^20 FLOPs 30B 600B
Tech Company 10^21 FLOPs 100B 2T
Large Tech 10^22 FLOPs 300B 6T
Google/Meta 10^23+ FLOPs 500B+ 10T+

Rule of thumb: D ≈ 20N always holds!

Key Takeaways

Scaling laws are predictable: Loss decreases predictably with N and D Chinchilla: Optimal is D ≈ 20N (balance parameters and data) GPT-3 was undertrained (300B tokens for 175B params) 70B model trained on 1.4T tokens beats 175B on 300B tokens Use scaling laws to forecast quality before training

-