Skip to content

Agent Framework Integration

Overview

LLMs are typically deployed as components within agent systems. This document bridges LLM Model deployment patterns with agent frameworks that use them.

-

When LLMs Are Used in Agents

LLMs power different layers in agent systems:

Agent System Architecture
 ↓
 - Perception Layer
 - Parse/understand user input
 - Reasoning Layer
 - **LLM** (choose actions)
 - Action Layer
 - Execute tools/skills
 - Learning Layer
 - Improve from feedback
 - Output Layer
 - Communicate results

The LLM is the "brain" - it decides what to do next.

-

Claude Agents API (Simplest Integration)

Location: 04 Claude Agents Api

LLM Considerations:

  • Model selection: Claude 3.5 Sonnet (fast+smart balance)
  • Context window: Use most of it (200K tokens)
  • Token management: Set max_tokens limit
  • Cost control: Track tokens_used from API

Deployment Pattern:

LLM Service (Claude API)
 ↓
Claude Agents SDK (handles tool calling)
 ↓
Your Application

Related LLM Topics:

-

LangGraph (Complex Workflow Integration)

Location: 06 Langgraph Framework

LLM Considerations:

  • Multiple LLM calls across nodes
  • Model selection by task complexity
  • State management for context
  • Token usage per node
  • Caching between nodes

Deployment Pattern:

LLM Service (Claude API / OpenAI)
 ↓
LangGraph StateGraph (workflow orchestration)
 ↓
Nodes process with LLM
 ↓
Tool execution
 ↓
Feedback loops

Related LLM Topics:

-

MCP Protocol (Tool Integration)

Location: 01 Mcp Protocol

LLM Considerations:

  • MCP doesn't change LLM behavior
  • LLM still processes and chooses tools
  • MCP standardizes tool interface
  • Reduces token waste on tool descriptions

Before MCP:

LLM must know about many tool formats
Each tool has different schema
More tokens wasted on tool descriptions

With MCP:

LLM knows one tool protocol (MCP)
All tools standardized via MCP server
Fewer tokens on redundant descriptions
Cleaner tool integration

Related LLM Topics:

-

Multi-Agent Systems (CrewAI, AutoGen)

Locations:

LLM Considerations:

  • Multiple LLM instances (one per agent)
  • LLM call amplification (more calls than single agent)
  • Context per agent vs shared
  • Communication overhead (LLM processing team messages)

Cost Impact:

Single agent: 10 calls/request, cost = $0.10
5-agent system: 50 calls/request, cost = $0.50

5x cost, but 2x capability → better ROI usually

Related LLM Topics:

  • [Parallel Inference](/01-modeling/04-production/01-deployment/(load-balancing-request-routing/)
  • Cost Scaling

-

Context Window Management in Agents

Key Decision: How much context to give LLM?

# Pattern 1
all_messages  LLM
Pros: Complete picture
Cons: Expensive, slow

# Pattern 2
last_20_messages  LLM
Pros: Balance cost/quality
Cons: Lose old context

# Pattern 3
summary_of_old + recent  LLM
Pros: Remember everything, lower cost
Cons: More complex implementation

Agent Framework Support:

  • Claude API: Manual (you decide)
  • LangGraph: Built-in via State
  • CrewAI: Automatic memory management

Related LLM Topics:

-

Scaling LLM Deployments for Agents

Pattern 1: Single LLM Service

All Agents → Claude API (shared)
Pros: Simple, consistent
Cons: Bottleneck under load

Pattern 2: Distributed LLM Services

Agent 1 → Claude API instance 1
Agent 2 → Claude API instance 2
Agent 3 → Claude API instance 3
Pros: Parallel, scalable
Cons: More complex

Pattern 3: Load-Balanced

Agents → Load Balancer → Claude API (auto-scaled)
Pros: Dynamic scaling
Cons: Complex setup

Related LLM Topics:

  • [Load Balancing](/01-modeling/04-production/01-deployment/(load-balancing-request-routing/)
  • Operations

LLM Model Selection in Agents

When to Use Different Models:

Model Best For Cost Speed Quality
Claude 3.5 Sonnet Most tasks Low Fast High
Claude 3 Opus Complex reasoning Medium Slower Very High
Smaller models Cheap tasks Very Low Very Fast Low-Medium

Agent Pattern:

class SmartAgentModelSelection:
 def select_model(self, task_complexity):
 if task_complexity < 0.3:
 return "claude-3-haiku" # Cheap
 elif task_complexity < 0.7:
 return "claude-3-5-sonnet" # Balanced
 else:
 return "claude-3-opus" # Best

Related LLM Topics:

-

Monitoring LLM Behavior in Agents

What to Monitor:

  1. Token Usage
Per request: Track tokens_used
Per agent: Aggregate by agent
Per framework: Compare LangGraph vs CrewAI
  1. Model Performance
Tool calling accuracy
Reasoning quality
Cost per successful action
  1. Inference Latency
Time from request to first token (TTFT)
Time to completion (latency)
Throughput (requests/sec)

Related LLM Topics:

-

Cross-Reference Map

LLM Modeling (This Book) Agent System Design (That Book)
───────────────────────────────────────────────────────────────
Architecture → Ch 3: Agent Architecture
 Ch 6: Tool Use

Inference → Ch 11: Frameworks
Optimization (How frameworks call LLMs)

Training/Fine-tuning → Ch 10: Advanced Topics
 (Self-evolution)

Production → Ch 9: Production Patterns
Deployment (How to deploy agents)

Operations → Ch 9: Production Patterns
Safety Ch 7: Safety & Reliability

Benchmarks → Ch 8: Evaluation
Performance (Agent benchmarking)

Best Practices Summary

  1. Choose Right Framework

  2. Simple: Claude Agents API

  3. Complex: LangGraph
  4. Team: CrewAI

  5. Optimize Tokens

  6. Set max_tokens

  7. Use context windows wisely
  8. Cache when possible

  9. Monitor Costs

  10. Track tokens per request

  11. Compare models by ROI
  12. Alert on cost spikes

  13. Scale Safely

  14. Load balance LLM calls

  15. Handle failures gracefully
  16. Test at higher load first

  17. Keep Updated

  18. Use latest models

  19. Monitor API changes
  20. Stay on top of security updates

-

Last Updated: August 9, 2026 Cross-Reference: Bridges LLM Modeling and Agent System Design knowledge bases