Skip to content

Monitoring & Observability: Production LLM Systems

Overview

Monitoring tracks system health (latency, throughput, errors). Observability enables understanding system behavior through logs, metrics, traces. Essential for production LLM systems to maintain quality and performance.

  • Metrics: Latency (P50/P99), throughput, error rates, GPU utilization
  • Observability: Distributed tracing, detailed logging, visualization
  • Challenges: Unique to LLMs (quality degradation, prompt injection)
  • Goal: Detect issues before users notice, maintain SLAs

Key Metrics

Latency Metrics

Latency breakdown:

Total latency = Queue time + Model time + Post-processing

Queue time:
  - How long request waits before processing starts
  - Should be <10% of total for good UX

Model time:
  - Forward pass, token generation
  - Usually 90%+ of total latency
  - Depends on model size, sequence length, batch size

Post-processing:
  - Response parsing, logging, etc.
  - Usually <5% of total

SLA targets (typical):

Model           P50     P99
──────────────────────────
7B model        100ms   500ms
70B model       500ms   2000ms (2s)
Mixture (small) 80ms    300ms

Monitoring:
  - Track P50 (median - typical user experience)
  - Track P99 (worst case - unhappy users)
  - Alert if P99 > threshold
  - P99 blowup often indicates queueing issue

Alert thresholds:
  - P99 > 2x normal → Investigate
  - P99 > 5x normal → Page on-call
  - P99 > 10x normal → Incident!

Throughput Metrics

Throughput: Tokens generated per second (GPU utilization metric)

Calculation:
  - Tokens/sec = Total tokens generated / Total time
  - Example: 1000 requests, 100 tokens each, 10 seconds
    - 100,000 tokens / 10 sec = 10K tokens/sec

Factors affecting throughput:

1. Model size
  - 7B model: 200-500 tokens/sec per GPU
  - 70B model: 30-80 tokens/sec per GPU
  - Larger model = lower throughput

2. Batch size
  - Batch 1: 100 tokens/sec
  - Batch 8: 300 tokens/sec (3x!)
  - Batch 32: 400 tokens/sec (diminishing returns)
  - More batching = higher throughput

3. Hardware
  - A100: 300 tokens/sec
  - H100: 600 tokens/sec (2x faster)
  - Better GPU = higher throughput

Target throughput:
  - Depends on model and deployment:
  - Research: 100+ tokens/sec acceptable
  - Production: 300+ tokens/sec (batch utilization)
  - High-scale: 1000+ tokens/sec (many concurrent users)

Alerting:
  - If throughput < baseline by 20% → investigate
  - Could indicate:
    - GPU throttling (temperature, power)
    - Resource contention (other processes)
    - Hardware degradation

Quality Metrics

Challenge specific to LLMs: Quality degradation

Sources of degradation:
1. Model drift (retraining affects quality)
2. Input distribution shift (new domains, prompt injection)
3. Numerical precision issues (FP16 rounding)
4. Outdated knowledge (training data old)

Monitoring quality:

A. Automatic metrics (continuous):
  - Self-consistency: Generate response 3x, check if consistent
    - Low consistency → potential quality issue
  - Perplexity on reference set: Keep benchmark, monitor over time
    - Increasing perplexity → model degrading
  - Token probability: Average log prob of generated tokens
  - Sharp drop → potential issue

B. Human metrics (periodic):
  - Sample responses weekly
  - Have humans rate quality
  - Track trend over time
  - Alert if trend negative

C. Proxy metrics:
  - Response length (unusual patterns)
  - Error token frequency
  - Repetition patterns
  - These correlate with quality issues

Example monitoring:
```python
def monitor_response_quality(response):
    """Continuous quality monitoring"""

    metrics = {}

    # Self-consistency: Generate 3 times
    responses = [generate(prompt) for _ in range(3)]

    # Check similarity
    similarity = compute_similarity(responses[0], responses[1:])
    metrics['consistency_score'] = similarity  # 0-1

    # Log probability
    log_probs = get_log_probs(response)
    metrics['avg_log_prob'] = sum(log_probs) / len(log_probs)

    # Repetition detection
    metrics['repetition_ratio'] = detect_repetition(response)

    # Alert if any metric unusual
    if metrics['consistency_score'] < 0.7:
        alert("Low consistency detected!")

    if metrics['avg_log_prob'] < -5:  # Very low confidence
        alert("Low confidence predictions!")

    return metrics
---

## Distributed Tracing

### Request-Level Observability
Challenge: Request goes through many components - Load balancer - Router - Model server - GPU - Response handler

Need to trace: Where does latency come from?

Distributed trace solution:

Request 1 - Trace ID: abc123 - Load Balancer (5ms) - Span ID: lb001 - Router (2ms) - Span ID: router001 - Model server (1200ms) - Span ID: server001 - GPU computation (1000ms) - Span ID: gpu001 - Batching overhead (100ms) - Span ID: batch001 - Output processing (100ms) - Span ID: output001 - Total: 1207ms

Analysis: - GPU took 1000ms (bottleneck, expected) - Batching took 100ms (could optimize) - Overhead ~200ms (within reasonable) - No major issues!

Implementation (OpenTelemetry):

from opentelemetry import trace
from opentelemetry.exporter.jaeger import JaegerExporter
from opentelemetry.sdk.trace import TracerProvider
from opentelemetry.sdk.trace.export import SimpleSpanProcessor

# Setup tracing
jaeger_exporter = JaegerExporter(agent_host_name="localhost", agent_port=6831)
trace.set_tracer_provider(TracerProvider())
trace.get_tracer_provider().add_span_processor(SimpleSpanProcessor(jaeger_exporter))

# Use in code
tracer = trace.get_tracer(__name__)

def process_request(request):
    with tracer.start_as_current_span("process_request") as span:
        span.set_attribute("request_id", request.id)

        with tracer.start_as_current_span("model_inference"):
            output = model(request)

        with tracer.start_as_current_span("output_processing"):
            response = format_response(output)

    return response
---

## Error Tracking

### Error Categories
Error types in LLM systems:

  1. Inference errors
  2. CUDA out of memory
  3. Model loading failed
  4. Invalid input
  5. GPU timeout

  6. Quality errors

  7. Gibberish output
  8. Hallucinations
  9. Inconsistent responses
  10. Prompt injection issues

  11. Performance errors

  12. Timeout (exceeds SLA)
  13. Resource exhaustion
  14. Cascading failures
  15. Deadlocks

  16. Data errors

  17. Missing context
  18. Encoding issues
  19. Format mismatches

Monitoring: - Track error rate per category - Alert on sudden spikes

### Error Alerting

```python
class ErrorMonitor:
    def __init__(self):
        self.error_counts = defaultdict(int)
        self.baseline_rates = {}  # Normal error rate per type

    def log_error(self, error_type, error_details):
        """Log error and check if abnormal"""

        self.error_counts[error_type] += 1

        # Get baseline for this error type
        baseline = self.baseline_rates.get(error_type, 0.01)  # 1% default

        # Current error rate (errors/sec)
        current_rate = self.error_counts[error_type] / self.time_window

        # Alert if 5x above baseline
        if current_rate > baseline * 5:
            alert(f"Abnormal error spike: {error_type}")
            notify_on_call()

        # Log for analysis
        log_to_datadog({
            'error_type': error_type,
            'details': error_details,
            'rate': current_rate
        })


Dashboards and Alerts

Key Dashboard Panels

LLM Production Dashboard:

Panel 1: Latency trends
  - P50, P95, P99 over time
  - Color: Green <SLA, Yellow near SLA, Red over SLA
  - Alert: P99 > threshold

Panel 2: Throughput
  - Tokens/sec over time
  - Target line (expected throughput)
  - Alert: Throughput drop >20%

Panel 3: Error rate
  - Errors per second
  - Breakdown by error type
  - Alert: Error rate spike

Panel 4: GPU utilization
  - GPU %, memory %
  - Per-GPU breakdown
  - Alert: Throttling, temp warnings

Panel 5: Quality metrics
  - Consistency scores
  - Confidence levels
  - Alert: Quality degradation

Panel 6: Cost tracking
  - Cost per query
  - Cost per token
  - Daily/monthly totals
  - Alert: Cost overages

Alert Thresholds

Metric                  Warning   Critical  Action
──────────────────────────────────────────────────
P99 Latency            2x normal  5x normal Page on-call
Throughput drop        10%        20%       Investigate
Error rate increase    2%         5%        Incident
GPU temperature        80°C       90°C      Reduce load
Memory OOM             First      Escalate  Immediate action
Quality score drop     5%         10%       Review model
Cost overrun           10% over   20% over  Alert finance

Alert routing:
  - Warning: Slack notification
  - Critical: Page on-call engineer
  - Incident: War room + leadership

Operational Best Practices

Canary Deployments

New model version deployment:

Stage 1: Canary (1% traffic)
  - Route 1% of requests to new model
  - Monitor: Quality, latency, errors
  - Duration: 1-2 hours
  - Success criteria: No quality regression, latency acceptable

Stage 2: Progressive (10% → 50% → 100%)
  - If canary OK, route 10%
  - Monitor for 30 minutes
  - Double traffic each stage
  - Continue until 100%

Stage 3: Rollback (instant if needed)
  - If any metrics bad: Rollback to previous
  - Automatic detection
  - No manual approval needed for automatic rollback

Result:
  - New model tested on real traffic
  - Fast rollback if issues
  - Confidence in deployments
  - Zero-downtime updates

Incident Response

LLM incident workflow:

1. Detection (automated alerts)
  - Metric threshold breached
  - Multiple metrics in alert state
  - Incident automatically created

2. Triage (on-call engineer)
  - Check dashboards
  - Look at logs
  - Determine severity
  - Notify stakeholders if critical

3. Investigation
  - Check recent changes
  - Look at distributed traces
  - Correlate with other systems
  - Identify root cause

4. Mitigation
  - Rollback if necessary
  - Scale resources
  - Implement quick fix
  - Restore service if impacted

5. Post-incident review
  - Document what happened
  - Why detection was late
  - How to prevent recurrence
  - Update runbooks

Key Takeaways

📊 Track P99 latency, not just average (outliers matter)
🔄 Distributed tracing: Understand where latency comes from
⚠️ Quality monitoring: Auto-detect degradation
🎯 Alert on anomalies, not just absolute thresholds
🚀 Canary deployments: Test new models safely