Skip to content

Cost Optimization Strategies

Overview

Cost optimization reduces both training costs (compute hours) and inference costs (per-query). Combining techniques achieves 10-100x cost reduction through hardware selection, parallelization, quantization, and efficiency techniques.

  • Training cost drivers: Compute hours, GPU hours, data transfer
  • Inference cost drivers: Token generation time, model size, concurrent users
  • Optimization: 10-100x cost reduction possible through technique combination
  • Trade-off: Quality vs. cost

-

Training Cost Optimization

Hardware Selection

GPU Cost Comparison (per hour):

GPU Cost/hr TFLOPS Efficiency (TFLOPS/$)
─────────────────────────────────────────────────
V100 $2.50 125 50
A100 $3.50 312 89
H100 $6.00 990 165
L40S $4.00 722 180

Recommendation by scale:
 - Small training (<100B FLOPs): A100 (best balance)
 - Medium training (100B-10T FLOPs): H100 (better efficiency)
 - Large training (>10T FLOPs): H100 (only practical option)

Cost example: Train LLaMA 7B

With A100 (8 GPUs):
 - Compute: 8 × $3.50/hr × 50 hours = $1,400
 - Data transfer: $100
 - Storage: $50
 - Total: $1,550

With H100 (8 GPUs):
 - Compute: 8 × $6/hr × 20 hours = $960 (4x faster)
 - Data transfer: $100
 - Storage: $50
 - Total: $1,110 (28% cheaper despite higher hourly cost!)

Key insight:
 - More expensive GPU/hour, but much faster
 - Amortizes fixed costs, saves money!

Parallelization Efficiency

Cost of parallelization:

Sequential training (impossible):
 - 1 GPU, many weeks → unfeasible

Data parallelism (efficient):
 - N GPUs, communication each step
 - Efficiency: 95% of ideal (only 5% waste)
 - Cost reduction: 18x for 16 GPUs

Pipeline parallelism (some overhead):
 - N GPUs, pipeline bubbles
 - Efficiency: 75% of ideal (25% waste)
 - Cost reduction: 12x for 16 GPUs
 - More complex setup

FSDP (balanced):
 - N GPUs, efficient communication
 - Efficiency: 80-90% (only 10-20% waste)
 - Cost reduction: 14x for 16 GPUs
 - Good balance of speed and complexity

Recommendation:
 - 8 GPUs: Use data parallelism
 - 16-64 GPUs: Use FSDP
 - 64+ GPUs: Consider pipeline + FSDP hybrid

Efficiency Techniques

Technique Speedup Cost Reduction
─────────────────────────────────────────────
Mixed precision (AMP) 2-3x 50-60%
Gradient checkpointing 1.3x 20% (memory, but compute cost up)
Flash Attention 1.5x 30%
Gradient accumulation 1.2x 10% (batching amortization)
All combined ~5x 80%

Cumulative effect:
Training cost with all techniques:
 - Baseline: $10,000 (100 GPU-hours @ $100/hr)
 - +AMP: $5,000 (2.5x faster)
 - +Flash Attention: $3,300 (1.5x more)
 - +FSDP (16 GPUs): $200 (best utilization!)
 - Total: 50x cost reduction!

Reality check:
 - Baseline: 100 GPU-hours
 - With optimization: 2-5 GPU-hours
 - Cost: $100-500 (vs $10,000!)

-

Inference Cost Optimization

Per-Query Costs

Cost breakdown per query:

Base cost: GPU time
 - Generate 100 tokens @ 100 tokens/sec = 1 second
 - GPU hour cost: $0.003-0.01 per second
 - Query cost: $0.003-0.01

Optimization: Use smaller model
 - 70B model: 0.3x throughput (3x slower)
 - 7B model: 0.9x baseline throughput
 - 7B query: $0.001-0.003 (1/3 of 70B)
 - Cost: 3-10x reduction!

Optimization: Use quantization
 - int4 quantization: 2x faster, 4x smaller
 - Same 70B, but 2x throughput
 - Query cost: $0.0015-0.005 (1/2 of FP16)
 - Cost: 2x reduction

Combined (7B + int4 quantization):
 - Base 70B: $0.01 per 100 tokens
 - 7B: $0.003
 - 7B + int4: $0.0015
 - Total: 6-10x cost reduction!

Batching and Utilization

Utilization impact on cost:

Scenario 1: Low utilization (sparse requests)
 - Requests come slowly
 - GPU often idle
 - Effective cost: High (paying for idle time)

Example:
 - 10 requests/hour
 - 1 GPU, $10/hour
 - Cost per request: $1.00

Scenario 2: High utilization (many requests)
 - Requests batched efficiently
 - GPU always busy
 - Effective cost: Low

Example:
 - 1000 requests/hour (same throughput, better batching)
 - 1 GPU, $10/hour
 - Cost per request: $0.01

Improvement: 100x cost reduction through batching!

Recommendation:
 - High volume → Amortize GPU cost
 - Batch many requests together
 - Share computation
 - 10-100x per-request cost savings

Strategy: Training vs Fine-tuning

Cost Comparison

Scenario: Need domain-specific LLM

Option 1: Train from scratch
 - Training cost: $100,000-1,000,000
 - Quality: Best possible
 - Time: Months

Option 2: Fine-tune pre-trained
 - Pretrained model: Free or $100-1000 (download)
 - Fine-tune cost: $100-10,000
 - Quality: 90-95% of option 1
 - Time: Days to weeks

Option 3: Use pre-trained + LoRA
 - Pretrained model: Free
 - LoRA training: $1-100
 - Quality: 85-90% of option 1
 - Time: Hours to days

Cost-benefit:
 - Option 1: Highest quality, highest cost
 - Option 2: Good quality, moderate cost (usually best)
 - Option 3: Acceptable quality, minimal cost
 - Recommendation: Option 2 or 3 (pre-train amortizes cost!)

Cost Sharing Through Adaptation

Multi-task model sharing:

Scenario: Need 5 domain-specific models
 - Option A: Train 5 models separately
 - Cost: 5 × $100,000 = $500,000
│
 - Option B: Train 1 base model, LoRA for 5 tasks
 - Base model: $100,000 (train once)
 - LoRA 1-5: $5 × $100 = $500
 - Total: $100,500
│
 - Option C: Use existing pre-trained, LoRA for 5 tasks
 - Pre-trained: Free
 - LoRA 1-5: $500
 - Total: $500

Cost reduction:
 - Option A: Baseline
 - Option B: 80% reduction (cost amortization)
 - Option C: 99.9% reduction (leverage pre-training)!

Key insight:
 - Pre-training cost >> fine-tuning
 - Massive ROI from reusing pre-trained models

-

Deployment Cost Optimization

Infrastructure Decisions

Options for deployment:

Option 1: Buy GPUs ($50K-100K upfront)
 - Initial: $100K
 - Running: $1-5K/month (power, cooling)
 - 3 years: $130-180K total
 - Good for: High traffic, long-term

Option 2: Cloud on-demand ($0.003-0.01/token)
 - Upfront: $0
 - Usage: Pay per query
 - 3 years (1B tokens): $3K-10K
 - Good for: Variable traffic, short-term, low volume

Option 3: Spot instances (AWS, GCP) ($0.001-0.003/token)
 - Upfront: $0
 - Usage: 70-80% discount vs on-demand
 - 3 years (1B tokens): $1K-3K
 - Good for: Can tolerate interruptions, cost-sensitive

Option 4: Serverless (Lambda, etc.)
 - Upfront: $0
 - Usage: $0.01-0.05 per invocation + compute
 - Good for: Occasional requests, extreme simplicity

Break-even analysis:
 - High volume (>10M tokens/month) → Buy GPUs
 - Medium volume (1-10M tokens) → Spot instances
 - Low volume (<1M tokens) → On-demand or serverless

Autoscaling for Cost

class CostOptimizedScaling:
 """Auto-scale based on cost, not just demand"""

 def should_add_gpu(self, queue_depth, cost_per_hour):
 """
 Add GPU only if cost-effective
 """
 # Current wait time
 wait_time = queue_depth / self.throughput

 # Cost of waiting (user churn, SLA penalties)
 cost_of_waiting = wait_time * self.penalty_per_second

 # Cost of adding GPU
 cost_of_gpu = cost_per_hour / 3600 # per second

 # Add GPU only if it pays for itself
 return cost_of_waiting > cost_of_gpu * 2 # 2x margin

# Result

-

Cost Monitoring & Optimization

Key Metrics

Metrics to track:

1. Cost per query
 - Total cost / number of queries
 - Goal: Minimize

2. Cost per token generated
 - Total cost / total tokens
 - Goal: Minimize (use for billing)

3. GPU utilization
 - % of time GPU is computing (vs idle)
 - Goal: >80% (higher is more cost-efficient)

4. Cost per SLA compliance
 - Cost to maintain latency target
 - Goal: Optimize trade-off

5. Model efficiency
 - Quality per unit compute cost
 - Goal: Maximize (quality/$)

Cost Optimization Roadmap

Phase 1: Reduce training cost (one-time)
 - Switch to H100 (better efficiency)
 - Use AMP + Flash Attention
 - Use FSDP for parallelization
 - Goal: 80% training cost reduction

Phase 2: Reduce inference cost (ongoing)
 - Quantize to int4
 - Use LoRA for personalization (not full retraining)
 - Batch requests aggressively
 - Goal: 60% inference cost reduction

Phase 3: Optimize deployment
 - Use spot instances for variable load
 - Auto-scale based on cost metrics
 - Monitor and optimize constantly
 - Goal: 50% deployment cost reduction

Total impact: ~90% cost reduction possible!

-

Key Takeaways

Pre-training cost >> fine-tuning cost (leverage pre-trained!) Hardware selection crucial (H100 faster despite higher $/hr) Parallelization efficiency matters (FSDP better than naive) Batching amortizes GPU cost (high utilization = low per-query cost) Combined techniques: 10-100x cost reduction possible

-

  • Distributed Training - Cost-efficient training strategies
  • Quantization - Inference cost reduction
  • Lora - Cost-efficient fine-tuning
  • Continuous Batching - Inference utilization
  • [Load Balancing & Request Routing](/01-modeling/04-production/01-deployment/(load-balancing-request-routing/) - Maximize throughput per GPU