Skip to content

Production Metrics

Overview

While research benchmarks measure capability, production metrics measure business value. A model can score 85% on MMLU but fail in production due to:

  • High latency (users wait > 10 seconds)
  • Expensive inference ($0.10 per request)
  • Unreliable quality (hallucinations in 15% of cases)
  • Poor scalability (can't handle 1000 concurrent users)

Production metrics align technical performance with business objectives.

-

Metric Categories

graph TD
 A["Production Metrics"] --> B["Performance Metrics"]
 A --> C["Quality Metrics"]
 A --> D["Economics Metrics"]
 A --> E["Reliability Metrics"]

 B --> B1["Latency<br/>P50, P99"]
 B --> B2["Throughput<br/>tokens/sec"]
 B --> B3["Availability<br/>Uptime %"]

 C --> C1["Accuracy<br/>Task-specific"]
 C --> C2["Hallucination Rate<br/>Factuality"]
 C --> C3["User Satisfaction<br/>NPS, Rating"]

 D --> D1["Cost per Inference<br/>$/completion"]
 D --> D2["Cost per User<br/>Monthly cost"]
 D --> D3["ROI<br/>Revenue impact"]

 E --> E1["Error Rate<br/>Failures"]
 E --> E2["Safety Score<br/>Harmful content"]
 E --> E3["Consistency<br/>Drift detection"]

-

1. Performance Metrics

Latency

Definition: Time from request to complete response

# Types of latency

# Time-to-First-Token (TTFT)
# When does generation start?
 - # ├─ Tokenization: 10ms
 - # ├─ Prompt processing: 100ms
 - # └─ First token: 20ms
# = TTFT

# Token-to-Token Latency
# Time between consecutive tokens
 - # ├─ Each token generation: 50ms
 - # ├─ Decoding overhead: 5ms
# = Per-token latency

# End-to-End Latency
# Full completion time
# = TTFT + (num_tokens * per_token_latency)
# = 130ms + (100 tokens * 55ms) = 5.6s

# Percentiles matter!
percentiles = {
 "P50": 2.1, # Median (50% of requests)
 "P95": 5.4, # 95th percentile (95% faster)
 "P99": 8.2, # 99th percentile (99% faster)
 "max": 45.0, # Outliers
}

# Interpretation:
# P99 = 8.2s means 1 in 100 users waits > 8.2 seconds
# SLA

Target Latencies:

Use Case P50 P99 Notes
──────────────────────────────────────────────────
Search Results < 50ms < 200ms Must be instant
Chat (interactive) < 500ms < 2s User waiting
Document Analysis < 5s < 30s Batch processing
Code Generation < 2s < 10s Developer workflow
Email Draft < 5s < 60s Background

Optimization Strategies:

# 1. Reduce Prompt Processing
# Use KV-cache to skip redundant computation
# Use attention masks to skip tokens

# 2. Reduce Token-Generation Latency
# Use Flash Attention (2x faster)
# Use Speculative Decoding (3x faster)
# Quantize model (1.5x faster, slight quality loss)

# 3. Reduce End-to-End Latency
# Use continuous batching (increase throughput, same latency)
# Cache embeddings for common prompts
# Use distilled models (7B instead of 70B)

# Typical improvements:
baseline = 2.1 # seconds P50
+ continuous_batching = 2.1 # (no latency change)
+ flash_attention = 1.0
+ speculative_decoding = 0.35
+ quantization = 0.25
= improved = 0.25 # seconds (8.4x faster!)

Throughput

Definition: How many tokens/requests can be processed per second

# Token Throughput (tokens/sec)
max_tokens_per_sec = 1000
avg_tokens_per_request = 200

# Request Throughput
requests_per_sec = max_tokens_per_sec / avg_tokens_per_request
# = 1000 / 200 = 5 requests/sec

# Batch Efficiency
# How much faster with batching?
single_request_latency = 2.0 # seconds
batch_4_latency = 2.2 # seconds (+10% for batching overhead)

throughput_single = 1 / single_request_latency = 0.5 req/sec
throughput_batch = 4 / batch_4_latency = 1.8 req/sec
improvement = 1.8 / 0.5 = 3.6x

# Batching yields 3.6x throughput improvement!

Capacity Planning:

# Goal
# Each user generates
# Average response

required_tokens_per_sec = (1000_users * 1 / 60) * 200
# = 16.67 * 200 = 3333 tokens/sec

# With A100 GPU
# Headroom

# Without batching (single A100 = 1000 tok/sec):
# 3333 / 1000 = 3.3 GPUs needed (expensive!)

# With batching
# Huge cost savings with batching

-

2. Quality Metrics

Accuracy (Task-Dependent)

Classification Tasks:

from sklearn.metrics import accuracy_score, f1_score, confusion_matrix

predictions = model.generate(test_prompts)
ground_truth = test_labels

# Metrics
accuracy = accuracy_score(ground_truth, predictions)
f1 = f1_score(ground_truth, predictions, average="weighted")
cm = confusion_matrix(ground_truth, predictions)

# Interpretation
accuracy = 0.92 # 92% correct classifications
f1 = 0.89 # Good precision and recall balance

# Track over time
metrics_history = {
 "2024-08-01": {"accuracy": 0.88, "f1": 0.85},
 "2024-08-08": {"accuracy": 0.92, "f1": 0.89},
 "2024-08-15": {"accuracy": 0.91, "f1": 0.88},
 # Slight degradation on Aug 15 (investigate!)
}

Generation Tasks (BLEU/ROUGE):

from rouge_score import rouge_scorer
from nltk.translate.bleu_score import sentence_bleu

# ROUGE (for summarization, translation)
scorer = rouge_scorer.RougeScorer(['rouge1', 'rougeL'])

reference = "The quick brown fox jumps"
generated = "The quick fox jumps over"

scores = scorer.score(reference, generated)
# ROUGE-1
# ROUGE-L

# BLEU (for translation)
reference = ["the", "quick", "brown", "fox"]
generated = ["the", "quick", "fox"]
bleu = sentence_bleu([reference], generated)
# BLEU

# Thresholds
ROUGE_L_threshold = 0.45 # Minimum quality

Semantic Similarity (for QA, RAG):

from sentence_transformers import util

# Compare generated vs expected
expected = "Paris is the capital of France"
generated = "France's capital is Paris"

embedding_model = SentenceTransformer('all-MiniLM-L6-v2')
embeddings = embedding_model.encode([expected, generated])

similarity = util.pytorch_cos_sim(embeddings[0], embeddings[1])
# similarity = 0.94 (nearly identical semantically)

# Threshold

Hallucination Rate

Definition: % of responses containing false/fabricated information

# Evaluation methodology
test_queries = [
 "What is the capital of France?",
 "When did World War 2 end?",
 "Who is the current president of the USA?"
]

predictions = model.generate(test_queries)

# Fact-checking (manual or automated)
fact_checker = AutomaticFactChecker() # Using knowledge base

results = {
 "query": "What is the capital of France?",
 "response": "Paris is the capital of France",
 "is_factual": True,
 "confidence": 0.98
}

# Calculate hallucination rate
hallucinations = sum(1 for r in results if not r['is_factual'])
hallucination_rate = hallucinations / len(results)
# = 0.05 (5% hallucination rate)

# Benchmark thresholds
thresholds = {
 "Excellent": "< 2%",
 "Good": "2-5%",
 "Acceptable": "5-10%",
 "Poor": "> 10%"
}

# Track hallucinations per domain
domain_hallucinations = {
 "Medical facts": 3.2, # Higher in specialized domains
 "General knowledge": 4.8,
 "Recent events": 8.1, # Highest (training data older)
 "Mathematics": 1.2 # Lowest (deterministic)
}

Reduction Strategies:

# 1. Retrieval-Augmented Generation (RAG)
# Provide facts from knowledge base
# Reduces hallucinations by 50-70%

# 2. Fact Verification
# Check outputs against knowledge base
# Flag unverified claims

# 3. Temperature Reduction
# temperature = 0.3 (more confident, deterministic)
# vs
# temperature = 0.7 (more creative, more hallucinations)

# 4. Fine-tuning on Factual Data
# Train on QA pairs with sources
# Improves factuality by 20-30%

User Satisfaction (NPS, Rating)

Net Promoter Score (NPS):

# Survey
# Scale

responses = [9, 8, 10, 7, 2, 1, 8, 9, 6, 7]

# Calculate NPS
promoters = len([r for r in responses if r >= 9]) # 9-10
detractors = len([r for r in responses if r <= 6]) # 0-6
passives = len([r for r in responses if 7 <= r <= 8]) # 7-8

nps = ((promoters - detractors) / len(responses)) * 100
# = ((4 - 2) / 10) * 100 = 20

# NPS Interpretation
nps_score = {
 "> 50": "Excellent",
 "0-50": "Good",
 "-50-0": "Problematic",
 "< -50": "Critical"
}

# For LLM products
typical_nps = {
 "ChatGPT": 65, # Excellent
 "Claude": 60,
 "In-house LLM": 35, # Needs improvement
 "Legacy system": 15
}

5-Star Rating System:

# After each interaction
rating_distribution = {
 5: 45, # 45% excellent
 4: 30, # 30% good
 3: 15, # 15% neutral
 2: 7, # 7% poor
 1: 3 # 3% terrible
}

average_rating = sum(k*v for k,v in rating_distribution.items()) / sum(rating_distribution.values())
# = 4.07 / 5.0 (Good)

# Track over time
rating_trends = {
 "Week 1": 4.1,
 "Week 2": 4.0,
 "Week 3": 3.9, # Declining (investigate!)
 "Week 4": 3.7
}

# Action
# Response

-

3. Economic Metrics

Cost per Inference

# Calculate total cost of ownership

# Fixed Costs (infrastructure)
gpu_cost_monthly = 2000 # A100 GPU

# Variable Costs (per request)
cost_per_1m_tokens = 5 # Your inference cost
tokens_per_request = 200

cost_per_request = (tokens_per_request / 1_000_000) * cost_per_1m_tokens
# = 0.000001 (one millionth of a dollar)

# Comparison
comparison = {
 "GPT-4 API": "$0.006 per 1K tokens → $0.0012 per request",
 "Claude API": "$0.003 per 1K tokens → $0.0006 per request",
 "Self-hosted": "$0.000005 per request (your cost)"
}

# Volume Break-Even
monthly_requests = 100_000

api_cost = monthly_requests * 0.0006 # Claude
self_hosted_cost = gpu_cost_monthly + (monthly_requests * 0.000005)

# = $60 (Claude) vs $2000.5 (self-hosted)
# → Use API for low volume

monthly_requests = 10_000_000
api_cost = monthly_requests * 0.0006 # = $6,000
self_hosted = gpu_cost_monthly + (10_000_000 * 0.000005) # = $52

# → Self-hosted wins for high volume

Cost per User

# How much does each user cost?

total_monthly_cost = 10_000 # Infrastructure + API
monthly_active_users = 5_000

cost_per_user = total_monthly_cost / monthly_active_users
# = $2 per user per month

# Unit Economics
revenue_per_user = 10 # Subscription: $10/month
profit_per_user = revenue_per_user - cost_per_user
# = $10 - $2 = $8 (80% gross margin)

# Revenue Breakdown
revenue = {
 "LLM inference": "$2k/month (20%)",
 "Infrastructure": "$3k/month (30%)",
 "Engineering": "$5k/month (50%)"
}

# Cost Reduction Strategy
cost_reduction = {
 "Quantize model": "-30% (inference)",
 "Distill to smaller model": "-50% (inference)",
 "Batch requests": "-20% (infrastructure)",
 "Implement caching": "-40% (API calls)"
}

ROI of Model Upgrades

# Should we upgrade from Model A to Model B?

# Current State (Model A
model_a = {
 "accuracy": 0.75,
 "latency": 100, # ms
 "cost": 1_000, # $/month
 "user_satisfaction": 3.8
}

# Option B (Model B
model_b = {
 "accuracy": 0.87,
 "latency": 500, # ms
 "cost": 5_000, # $/month
 "user_satisfaction": 4.3
}

# Impact Analysis
# Accuracy improvement
# → Expected

# Latency increase
# → Risk

# Cost increase
# → Impact

# Decision Matrix
roi_analysis = {
 "revenue_impact": 1.10, # 10% increase
 "churn_impact": 0.93, # 7% churn
 "cost_impact": 5.0, # 5x cost increase

 "net_impact": 0.93 * 1.10 / 5.0
 # = 0.20 (only 20% of revenue benefit after churn and cost)

 "verdict": "Upgrade only if churn factor is < 2%"
}

-

4. Reliability Metrics

Error Rate & Availability

# Uptime
total_time = 30 * 24 * 60 # 30 days in minutes
downtime = 45 # 45 minutes of downtime

uptime_percentage = (total_time - downtime) / total_time * 100
# = (43200 - 45) / 43200 * 100 = 99.89% (4 nines)

# SLA (Service Level Agreement)
sla_targets = {
 "Standard": 99.5, # 3.6 minutes downtime/month
 "High Availability": 99.9, # 43 seconds downtime/month
 "Critical": 99.99, # 4 seconds downtime/month
}

# Error Rate
total_requests = 1_000_000
failed_requests = 2_500

error_rate = failed_requests / total_requests
# = 0.25% (0.0025)

# Error types
error_breakdown = {
 "Timeout": 40, # 1000 requests
 "Out of Memory": 600,
 "Inference Failure": 500,
 "Invalid Input": 400
}

# Action

Model Degradation Detection

# Track performance over time to detect drift

daily_metrics = {
 "2024-08-01": {"accuracy": 0.92, "latency": 150, "errors": 0.15},
 "2024-08-02": {"accuracy": 0.92, "latency": 152, "errors": 0.16},
 "2024-08-05": {"accuracy": 0.88, "latency": 180, "errors": 0.25},
 # ↑ Significant degradation detected
}

# Alert thresholds
thresholds = {
 "accuracy_drop": 0.03, # Alert if drop > 3 points
 "latency_increase": 1.2, # Alert if > 20% increase
 "error_rate_increase": 0.05 # Alert if > 5%
}

# Automatic Monitoring
@cron_daily
def check_model_health():
 current = get_latest_metrics()
 historical = get_historical_metrics()

 for metric, threshold in thresholds.items():
 if abs(current[metric] - historical[metric]) > threshold:
 alert(f"Model degradation detected: {metric}")
 log_incident()
 trigger_rollback()

# Root Causes (investigate degradation)
possible_causes = [
 "Data distribution shift (new types of queries)",
 "Model fine-tuning hurt base model",
 "Infrastructure change (different GPU)",
 "Cache invalidation",
 "Prompt template change",
 "Upstream API change (RAG knowledge base)"
]

-

Dashboard Setup

Recommended Metrics Dashboard:

dashboard_metrics = {
 "Performance": [
 "P50 Latency (target: < 500ms)",
 "P99 Latency (target: < 2s)",
 "Throughput (tokens/sec)",
 "Availability % (target: 99.9)"
],

 "Quality": [
 "Accuracy (task-specific)",
 "Hallucination Rate (target: < 5%)",
 "User Rating (target: > 4.0)",
 "NPS Score (target: > 50)"
],

 "Economics": [
 "Cost per Request",
 "Cost per User",
 "Revenue per User",
 "Gross Margin %"
],

 "Operations": [
 "Error Rate (target: < 0.5%)",
 "Model Degradation (trend)",
 "API Usage (quota tracking)",
 "GPU Utilization %"
]
}

# Visualization
# - Real-time dashboard (Grafana, Datadog)
# - Daily reports (email)
# - Alerts (Slack) when thresholds crossed
# - Weekly deep-dive reviews

-

Real-World Example: LLM Inference Service

# Hypothetical service

# Targets
targets = {
 "P99 Latency": "< 5 seconds",
 "Throughput": "> 100 requests/hour",
 "Accuracy": "> 80% (ROUGE-L)",
 "Hallucination": "< 3%",
 "Cost per Request": "< $0.01",
 "Uptime": "> 99.9%",
 "User Satisfaction": "> 4.0/5.0"
}

# Current Performance (Week 1)
week1 = {
 "P99 Latency": 8.5, # Above target (5s)
 "Throughput": 80, # Below target (100)
 "Accuracy": 78, # Below target (80)
 "Hallucination": 4.2, # Above target (3)
 "Cost per Request": 0.015, # Above target ($0.01)
 "Uptime": 99.2, # Close
 "User Satisfaction": 3.2 # Below target
}

# Optimization Plan
actions = {
 "Latency": "Implement Flash Attention (-40%)",
 "Throughput": "Enable continuous batching (+30%)",
 "Accuracy": "Fine-tune on domain data (+5%)",
 "Hallucination": "Add RAG + fact verification",
 "Cost": "Switch to cheaper inference framework",
 "Satisfaction": "Improve output quality"
}

# Results (Week 4 after optimization)
week4 = {
 "P99 Latency": 4.2, # Meets target
 "Throughput": 140, # Exceeds target
 "Accuracy": 83, # Meets target
 "Hallucination": 2.1, # Below target
 "Cost per Request": 0.007, # Below target
 "Uptime": 99.95, # Excellent
 "User Satisfaction": 4.3 # Exceeds target
}

# Success

-

References

Monitoring Tools

  • Prometheus + Grafana: Open-source metrics + visualization
  • DataDog: Commercial monitoring platform
  • New Relic: APM + infrastructure monitoring
  • CloudWatch: AWS native monitoring
  • Weights & Biases: ML-specific experiment tracking

Papers

-

Last Updated: 2026-08-09