Part 9¶
Overview¶
Getting an agent to work in a lab is one thing. Getting it to reliably work in production at scale is another.
This section covers the patterns that separate prototype agents from production systems:
- How to keep agent state consistent across failures
- When to route work to different agents
- How to let humans oversee critical decisions
- How to manage expensive context windows
- How to see what's happening in production
- How to recover gracefully from failures
Key Insight: Production is about tradeoffs. Safety vs speed, cost vs quality, autonomy vs control.
Chapter Statistics¶
| Metric | Value |
|---|---|
| Topic Files | 6 comprehensive guides |
| Total Words | 11,500+ |
| Code Examples | 55+ production-grade |
| Deployment Patterns | 20+ real patterns |
| Warnings | 18+ anti-patterns |
| Real-World Cases | 8+ case studies |
Production Reality¶
Lab Environment:
Controlled data
Unlimited time
High latency acceptable
Few requests
No failures expected
Production Environment:
Messy real-world data
Strict latency budgets
Millions of requests
Failures happen daily
Users hate delays
Key Production Concerns¶
| Concern | Impact | Pattern |
|---|---|---|
| Agent crashes | Lost context | State Management & Recovery |
| Can't see issues | Discover problems from users | Observability & Monitoring |
| Bad decisions | Regulatory/compliance risk | Human-in-the-Loop |
| Too many requests | Cascading failures | Routing & Load Balancing |
| Errors cascade | System outage | Error Recovery |
| Context too expensive | Token costs spiral | Context Management |
| Resource leaks | OOM, crashes | Resource Management |
-
Complete Chapter Organization¶
1. State Management & Recovery (2,100 words)¶
- Agent state and consistency
- Checkpointing strategies
- Recovery from failures
- Distributed state coordination
- Saga pattern for multi-step workflows
2. Routing & Escalation (1,800 words)¶
- Request routing strategies
- Load balancing
- Escalation policies
- Risk-based routing
- Queue management
3. Human-in-the-Loop (1,900 words)¶
- When to escalate to humans
- Approval workflows
- Feedback integration
- Hybrid autonomy model
- Scalable human oversight
4. Context Management (1,700 words)¶
- Context window economics
- Compression strategies
- Retrieval-augmented generation
- Context window optimization
- Token budgeting
5. Observability & Monitoring (1,900 words)¶
- Structured logging
- Trace collection
- Metrics and dashboards
- Alerting strategies
- Post-incident analysis
6. Error Recovery (2,100 words)¶
- Error classification
- Recovery strategies
- Circuit breakers
- Graceful degradation
- Fallback patterns
-
Learning Paths¶
Path 1: Full Production Deployment (6 hours)¶
- State Management - Keep state consistent
- Routing - Handle traffic
- Human Loop - Add oversight
- Context - Control costs
- Observability - See what's happening
- Error Recovery - Handle failures
Path 2: Quick to Production (3.5 hours)¶
- State Management - Stability first
- Observability - Visibility
- Error Recovery - Reliability
- Human Loop - Safety
Path 3: Scale Existing System (4 hours)¶
- Routing - Handle load
- Context - Reduce costs
- Error Recovery - Prevent cascades
- Observability - Track performance
Path 4: Optimize Running System (3 hours)¶
- Context - Reduce token costs
- Routing - Better load distribution
- Observability - Find bottlenecks
The Hybrid Autonomy Model¶
2025-2026 Best Practice:
Risk Level Autonomy Model Examples
─────────────────────────────────────────────────────────
Low (< $10) Fully Autonomous Standard support reply
→ Agent decides alone Simple data lookup
Medium Approve then Execute Refund > $50
($10-$100) → Human reviews Account modification
→ Agent executes Policy adjustment
High 👤 Human Decision Delete customer data
(> $100) → Agent recommends Override policy
→ Human decides High-value transaction
→ Agent executes
Critical 🚨 Manual with Review Payment processing
(> $10k) → Agent performs recon Regulatory decisions
→ Multiple humans review Security incidents
→ Human authorizes
-
Production Architecture¶
Request
↓
Input Validation (Safety)
↓
Routing (Cost/Capability)
- → Simple Agent (Low-risk)
- → Complex Agent (Medium-risk)
- → Human with Agent Assistance (High-risk)
↓
State Checkpoint (Durability)
↓
Execution with Error Handling
↓
Context Management (Cost Control)
↓
Monitoring & Observability
↓
Output Validation & Filtering
↓
State Persistence
↓
Response + Audit Trail
Common Production Patterns¶
Pattern 1: Simple Direct Execution (High Risk)¶
Request → Agent → Response
Problems:
- No error recovery
- State lost on crash
- Can't rollback
- No visibility
Pattern 2: Staged Execution (Recommended)¶
Request → Validate → Route → State Checkpoint → Execute →
Monitor → Validate Output → Persist → Response
Advantages:
- Each stage independent
- Can add safety checks
- Easy to observe
- Can rollback
Critical Warnings Summary¶
Production Mistakes:
- No state persistence (lost work on crash)
- No human oversight (risky decisions)
- No routing strategy (overwhelming single agent)
- Unbounded context (spiraling token costs)
- No observability (discover issues from users)
- No error recovery (cascading failures)
- No rate limiting (DDoS yourself)
Quick Start: Minimal Production Setup¶
class MinimalProductionAgent:
def __init__(self):
self.state_manager = StateManager()
self.router = Router()
self.monitor = Monitor()
self.error_handler = ErrorHandler()
def handle_request(self, request):
try:
# 1. Save state before processing
checkpoint = self.state_manager.checkpoint()
# 2. Route to appropriate handler
handler = self.router.select_handler(request)
# 3. Execute
result = handler.execute(request)
# 4. Monitor and validate
self.monitor.track(result)
# 5. Persist state
self.state_manager.persist(result)
return result
except Exception as e:
# Recover from checkpoint
self.error_handler.recover(checkpoint, e)
raise
-
Production Readiness Checklist¶
Before deploying agents to production:
- [] State Management: Checkpoint/recovery in place
- [] Routing: Request distribution strategy defined
- [] Human Loop: Escalation policy for high-risk
- [] Context: Token budgeting and compression
- [] Observability: Logging and metrics set up
- [] Error Handling: Recovery paths for all failures
- [] Monitoring: Dashboards and alerts configured
- [] Rate Limiting: Capacity protection in place
- [] Incident Response: On-call process documented
- [] Rollback Plan: How to revert changes quickly
Related Chapters¶
- Safety & Reliability (Ch 7): Safe execution
- Evaluation (Ch 8): Production metrics
- Tool Use (Ch 6): Tool observability
- Memory Systems (Ch 4): State persistence
- Planning & Reasoning (Ch 5): Decision tracing
Key Insights¶
- State is everything - Agents without persisted state are unreliable
- Humans at the helm - Critical decisions need human oversight
- Observability is non-negotiable - If you can't see it, you can't fix it
- Context matters - Token costs can spiral; manage proactively
- Failures will happen - Design for graceful degradation
- Routing for efficiency - Different requests need different agents
- Scale is a feature - Production must handle 10-1000x load
-
Start Reading¶
First time? → Start with State Management
Need reliability? → Start with Error Recovery
Need visibility? → Start with Observability
Handling high-risk decisions? → Start with Human In The Loop
Reducing costs? → Start with Context Management
Scaling up? → Start with Routing & Escalation
-
Last Updated: August 9, 2026 Status: Complete chapter guide (6 comprehensive topic files, 11,500+ words)