Graph-Based Orchestration¶
Overview¶
Graph-based orchestration represents agent workflows as directed graphs (DAGs), where nodes are tasks and edges are dependencies.
This is the approach used by production frameworks like LangGraph and is becoming the standard for complex agent systems.
Why Graphs?¶
Graphs naturally represent:
- Tasks (nodes)
- Dependencies (edges)
- Parallel work (independent paths)
- Conditional routing (branching)
- State flow (data between nodes)
Traditional thinking:
"Agent A then Agent B then Agent C"
Graph thinking:
A → B → C (sequential)
A → B (parallel)
→ C
A → {B or C} (conditional)
```
---
## Core Concepts
### Nodes
Represent units of work (agents, functions, decisions)
```python
class GraphNode:
name: str # "research"
function: Callable # what to execute
inputs: List # what it needs
outputs: List # what it produces
def execute(self, state):
return self.function(state)
```
### Edges
Represent dependencies and data flow
```python
# Simple edge
graph.add_edge("research", "analyze")
# Conditional edge
graph.add_conditional_edge(
"classify",
route_function, # Decides B or C
{"technical": "fix", "general": "help"}
)
# Multiple outputs
graph.add_edge("search", "analyze_1")
graph.add_edge("search", "analyze_2")
```
### State
Flows through the graph, transformed by each node
```python
class GraphState(TypedDict):
goal: str
query: str
research_results: List
analysis: Dict
final_output: str
```
---
## The Execution Model
```
Start State
↓
- ┌───────────────────┐
- Execute Ready │ (Find nodes with all inputs ready)
- Nodes │
- ┬───────────┘
│
- ┌───▼───────────┐
- Update State │ (Output becomes input to next)
- ┬───────────┘
│
- ┌───▼───────────┐
- Find Next │ (Which nodes are ready now?)
- Ready Nodes │
- ┬───────────┘
│
- ┌───▼───────────┐
- More Work? │
- ┬───────┬───┘
│ │
No Yes
│ │
- → Loop back
│
▼
Final State
```
---
## Example 1: Simple Research Pipeline
```python
from langgraph.graph import StateGraph, START, END
class ResearchState(TypedDict):
goal: str
papers: List[str]
summaries: Dict
report: str
# Create graph
graph = StateGraph(ResearchState)
# Add nodes (the work)
graph.add_node("search", search_papers)
graph.add_node("fetch", fetch_full_texts)
graph.add_node("analyze", analyze_papers)
graph.add_node("write", write_report)
# Add edges (the flow)
graph.add_edge(START, "search")
graph.add_edge("search", "fetch")
graph.add_edge("fetch", "analyze")
graph.add_edge("analyze", "write")
graph.add_edge("write", END)
# Compile and run
agent = graph.compile()
result = agent.invoke({"goal": "quantum computing"})
```
**Visual**:
```
START → Search → Fetch → Analyze → Write → END
```
---
## Example 2: Conditional Routing
```python
class TaskState(TypedDict):
task: str
task_type: str
result: str
graph = StateGraph(TaskState)
def classify_task(state):
"""Determine task type"""
classifier = LLM()
return {
"task_type": classifier.predict(state["task"])
}
def handle_research(state):
"""Handle research task"""
return {"result": "research done"}
def handle_code(state):
"""Handle coding task"""
return {"result": "code written"}
def handle_general(state):
"""Handle general task"""
return {"result": "response generated"}
# Nodes
graph.add_node("classify", classify_task)
graph.add_node("research", handle_research)
graph.add_node("code", handle_code)
graph.add_node("general", handle_general)
# Conditional edge
def route(state):
return state["task_type"]
graph.add_edge(START, "classify")
graph.add_conditional_edges(
"classify",
route,
{
"research": "research",
"coding": "code",
"general": "general"
}
)
graph.add_edge("research", END)
graph.add_edge("code", END)
graph.add_edge("general", END)
# Run
agent = graph.compile()
result = agent.invoke({"task": "Write code to sort array"})
```
**Visual**:
```
- START → Classify → ─ If research → Research → END
- If coding → Code → END
- If general → General → END
```
---
## Example 3: Parallel Branches
```python
class AnalysisState(TypedDict):
data: str
stats: Dict
patterns: List
summary: str
graph = StateGraph(AnalysisState)
def statistical_analysis(state):
return {"stats": compute_stats(state["data"])}
def pattern_detection(state):
return {"patterns": find_patterns(state["data"])}
def combine_results(state):
"""Combine parallel results"""
return {
"summary": f"Stats: {state['stats']}, Patterns: {state['patterns']}"
}
# Nodes
graph.add_node("stats", statistical_analysis)
graph.add_node("patterns", pattern_detection)
graph.add_node("combine", combine_results)
# Edges
graph.add_edge(START, "stats")
graph.add_edge(START, "patterns")
# Converge
graph.add_edge("stats", "combine")
graph.add_edge("patterns", "combine")
graph.add_edge("combine", END)
# Run (stats and patterns execute in parallel)
agent = graph.compile()
result = agent.invoke({"data": "..."})
```
**Visual**:
```
- ┌→ Stats ───┐
- START ──┤ └→ Combine → END
- → Patterns ┘
```
---
## Example 4: Complex Orchestration
```
Research Team Project:
```
```
Project Start
│
- ┌───────┼───────┐
│ │ │
Research Analysis Writing
│ │ │
- ┬───┼───┬───┘
│ │ │
Aggregation
│
Final Report
```
**Implementation**:
```python
class ProjectState(TypedDict):
goal: str
research_data: List
analysis: Dict
draft: str
final_report: str
graph = StateGraph(ProjectState)
# Add nodes
graph.add_node("research", research_agent.run)
graph.add_node("analyze", analysis_agent.run)
graph.add_node("write", writer_agent.run)
graph.add_node("aggregate", aggregate_results)
graph.add_node("finalize", finalize_report)
# Add edges
graph.add_edge(START, "research")
graph.add_edge(START, "analyze")
graph.add_edge(START, "write")
# Convergence point
graph.add_edge("research", "aggregate")
graph.add_edge("analyze", "aggregate")
graph.add_edge("write", "aggregate")
graph.add_edge("aggregate", "finalize")
graph.add_edge("finalize", END)
# Run (research, analyze, write execute in parallel)
agent = graph.compile()
```
---
## Advanced: Cycles and Loops
Some workflows need to loop:
```python
def should_revise(state):
"""Check if revision needed"""
quality = evaluate_output(state["output"])
return "revise" if quality < 0.8 else "finalize"
graph.add_conditional_edges(
"generate",
should_revise,
{
"revise": "generate", # Loop back!
"finalize": END
}
)
```
**Visual**:
```
START → Generate → Should Revise?
│ │
Yes No
│ │
- ┌─────────┘ │
- ▼
- → Generate END
(loop)
```
---
## State Management in Graphs
Each node receives state and returns updated state:
```python
def research_node(state: ResearchState) -> ResearchState:
# Receive state
goal = state["goal"]
existing = state.get("existing_papers", [])
# Do work
new_papers = search(goal)
# Return updated state
return {
**state, # Keep existing
"papers": existing + new_papers, # Add new
"search_complete": True
}
```
**Important**: States flow through the graph, accumulating information
-
## Comparison: Graph vs Procedural
### Procedural (Traditional)
```python
def workflow(goal):
papers = search(goal)
texts = fetch(papers)
analysis = analyze(texts)
report = write(analysis)
return report
```
**Pros**:
- Simple to read
- Easy to debug
- Linear flow
**Cons**:
- Can't handle parallelization
- Can't handle complex routing
- State management implicit
- Can't pause/resume
### Graph-Based (LangGraph)
```python
graph = StateGraph(...)
graph.add_node("search", search)
graph.add_node("fetch", fetch)
graph.add_node("analyze", analyze)
graph.add_node("write", write)
graph.add_edge(START, "search")
graph.add_edge("search", "fetch")
graph.add_edge("fetch", "analyze")
graph.add_edge("analyze", "write")
graph.add_edge("write", END)
agent = graph.compile()
```
**Pros**:
- Parallelization possible
- Complex routing easy
- State explicit
- Can checkpoint/resume
- Visualization support
**Cons**:
- More verbose
- Learning curve
- Slight overhead
-
## Production Graph Design Patterns
### Pattern 1: Sequential Pipeline
```
Search → Fetch → Analyze → Write → END
```
**Use for**: Multi-step workflows
**Framework**: LangGraph perfect fit
### Pattern 2: Fan-Out/Fan-In
```
- ┌→ A ──┐
- START┤→ B ──┼→ Merge
- → C ──┘
```
**Use for**: Parallel independent tasks
**Framework**: LangGraph multi-edge support
### Pattern 3: Conditional Routes
```
- ┌→ Path A → END
START→ Router
- → Path B → END
- → Path C → END
```
**Use for**: Task specialization
**Framework**: LangGraph conditional edges
### Pattern 4: Retry Loop
```
- ┌─ Good?
│ │
No Yes
│ │
- → Retry ──→ END
```
**Use for**: Reliability
**Framework**: LangGraph loops
---
## Graph Execution Strategies
### Strategy 1: Depth-First
Execute deepest path first
### Strategy 2: Breadth-First
Execute all at same depth before proceeding
### Strategy 3: Greedy
Execute ready nodes as soon as dependencies met
**Default**: Greedy (most efficient)
---
## Monitoring Graph Execution
```python
agent = graph.compile()
# Stream execution step-by-step
for step in agent.stream({"goal": "..."}):
print(f"Step: {step}")
# Output: {"search": {"papers": [...]}}
# Output: {"fetch": {"texts": [...]}}
# Output: {"analyze": {"analysis": {...}}}
# Trace execution
result = agent.invoke(
{"goal": "..."},
config={"callbacks": [tracer]}
)
# Check state at each step
for node_name, state in execution_trace:
print(f"{node_name}: {state}")
```
---
## Debugging Graphs
```python
# Visualize graph structure
print(graph.get_graph().draw_mermaid())
# Output:
# graph LR
# START --> search
# search --> fetch
# fetch --> analyze
# analyze --> write
# write --> END
# Debug state at node
def debug_node(state):
print(f"State before: {state}")
result = process(state)
print(f"State after: {result}")
return result
graph.add_node("debug_point", debug_node)
# Step through execution
import pdb
def breakpoint_node(state):
pdb.set_trace()
return state
```
---
## Best Practices for Graph Design
### 1. Keep Nodes Focused
```python
# Good
def research_node(state):
return {"papers": search(state["goal"])}
# Bad
def do_everything(state):
papers = search(state["goal"])
texts = fetch(papers)
analysis = analyze(texts)
return {"everything": [papers, texts, analysis]}
```
### 2. Make State Explicit
```python
# Good
class State(TypedDict):
goal: str
papers: List
analysis: Dict
# Bad
def process(state):
return {"data": "something"} # What is this?
```
### 3. Handle Errors
```python
# Good
def safe_node(state):
try:
return {"result": process(state)}
except ProcessError:
return {"result": None, "error": "Failed"}
# Bad
def unsafe_node(state):
return {"result": process(state)} # Throws if fails
```
### 4. Use Checkpointing
```python
# Save state at key points
memory = MemorySaver()
agent = graph.compile(checkpointer=memory)
# Resume from checkpoint
result = agent.invoke(
{"goal": "..."},
config={"configurable": {"thread_id": "session_123"}}
)
```
---
## Key Takeaways
1. **Graphs represent workflows naturally** - Nodes = tasks, edges = dependencies
2. **Parallelization built-in** - No special code needed
3. **State flows through graph** - Accumulated at each node
4. **Conditional routing simple** - Branch based on state
5. **Production frameworks support it** - LangGraph is standard
6. **Debuggable and monitorable** - Can trace execution
7. **Scalable** - From simple to complex
-
## Next Steps
- [Read Hierarchical Agents](/01-agent-design/03-architecture/05-hierarchical-agents/) - Manager-worker systems
- [Go To Pattern Selection](/01-agent-design/02-core-design-patterns/04-pattern-selection-framework/) - Choose your patterns
-
**Last Updated**: August 9, 2026