Coordination Strategies¶
Overview¶
When you have multiple agents, they need to coordinate. How do they communicate? How do they share information? How do they make decisions together?
This file covers the proven strategies for agent coordination in production systems (2025-2026).
The Coordination Problem¶
Multiple Agents:
Agent A Agent B
↓ ↓
- Must coordinate ─→
- ← Share results ←─
Challenges:
• Information sharing
• Decision making
• Conflict resolution
• Fault handling
• Consistency
```
-
## 4 Core Coordination Strategies
### Strategy 1: Shared State (Centralized)
**How it works**: All agents read/write to shared database
```python
class SharedStateCoordination:
def __init__(self):
self.shared_state = Database()
def agent_a_work(self):
# Read state
data = self.shared_state.get("research_data")
# Process
processed = self.process(data)
# Write back
self.shared_state.put("processed_data", processed)
def agent_b_work(self):
# Read from agent A's output
processed = self.shared_state.get("processed_data")
# Use it
analysis = self.analyze(processed)
# Write result
self.shared_state.put("analysis", analysis)
```
**Pros**:
- Simple to implement
- Guaranteed consistency
- Easy to debug
- Natural for sequential work
**Cons**:
- Central point of failure
- Scalability bottleneck
- Locking issues (who modifies what?)
- Not good for distributed systems
**When to use**:
- Small teams (2-5 agents)
- Sequential workflows
- Single-machine deployment
- Consistency critical
**Production systems**: 40% use this
---
### Strategy 2: Message Passing (Asynchronous)
**How it works**: Agents send messages to each other via queue
```python
class MessagePassingCoordination:
def __init__(self):
self.message_queue = MessageQueue()
def agent_a_work(self):
# Do work
result = self.research(query)
# Send message to Agent B
self.message_queue.send(
to="agent_b",
message={"type": "research_complete", "data": result}
)
def agent_b_work(self):
# Wait for message
message = self.message_queue.receive(from_agent="agent_a")
if message.type == "research_complete":
# Process Agent A's result
analysis = self.analyze(message.data)
```
**Pros**:
- Decoupled (agents independent)
- Scalable (add agents easily)
- Distributed-friendly
- Asynchronous (responsive)
**Cons**:
- Eventual consistency (delays)
- Harder to debug
- Message ordering issues
- Lost message handling
**When to use**:
- Large teams (5-20 agents)
- Parallel work
- Distributed systems
- Real-time requirements
**Production systems**: 35% use this
-
### Strategy 3: Event-Driven (Pub-Sub)
**How it works**: Agents publish events, others subscribe
```python
class EventDrivenCoordination:
def __init__(self):
self.event_bus = EventBus()
def agent_a_work(self):
# Do work
research_results = self.research()
# Publish event
self.event_bus.publish(
event_type="research_complete",
data=research_results
)
def agent_b_setup(self):
# Subscribe to events
self.event_bus.subscribe(
event_type="research_complete",
handler=self.on_research_complete
)
def on_research_complete(self, event):
# Handle event
analysis = self.analyze(event.data)
# Publish new event
self.event_bus.publish(
event_type="analysis_complete",
data=analysis
)
```
**Pros**:
- Loosely coupled (many-to-many)
- Scalable (add subscribers easily)
- Reactive (respond to events)
- Good for complex workflows
**Cons**:
- Ordering/causality issues
- Debugging complex flows
- Resource overhead
- Cascading failures
**When to use**:
- Complex workflows
- Many agents (10-20+)
- Real-time systems
- Need loose coupling
**Production systems**: 20% use this
---
### Strategy 4: Hierarchical (Manager-Worker)
**How it works**: One manager orchestrates workers
```python
class HierarchicalCoordination:
def __init__(self):
self.manager = ManagerAgent()
self.workers = {
"researcher": ResearcherAgent(),
"analyst": AnalystAgent(),
"writer": WriterAgent()
}
def run(self, goal):
# Manager decomposes goal
tasks = self.manager.decompose(goal)
# {
# "research": "Find papers on topic X",
# "analysis": "Analyze findings",
# "writing": "Write report"
# }
results = {}
# Manager assigns to workers
for task_name, task_desc in tasks.items():
worker = self.workers[task_name]
results[task_name] = worker.execute(task_desc)
# Manager aggregates
final = self.manager.aggregate(results)
return final
```
**Pros**:
- Clear control flow
- Scalable vertically (add layers)
- Easy to understand
- Good for large projects
**Cons**:
- Manager becomes bottleneck
- Manager complexity grows
- Not good for peer collaboration
- Single point of failure (manager)
**When to use**:
- Large, complex projects
- Clear task hierarchy
- Supervised workers
- Enterprise systems
**Production systems**: 30% use this
---
## Coordination Patterns
### Pattern 1: Sequential (Pipeline)
```
Agent A → Agent B → Agent C → Result
Agent A: Search for papers
Agent B: Fetch full texts
Agent C: Analyze and summarize
```
**Implementation**:
```python
def sequential_pipeline(goal):
result = goal
for agent in [search_agent, fetch_agent, analyze_agent]:
result = agent.run(result)
return result
```
**Characteristics**:
- Simple, predictable
- Slow (sequential)
- Good error handling (stop at failure)
- Used in 50% of workflows
---
### Pattern 2: Parallel (Map-Reduce)
```mermaid
graph TD
A["Input"] --> B["Agent A"]
A --> C["Agent B"]
A --> D["Agent C"]
B --> E["Merge Results"]
C --> E
D --> E
```
**Implementation**:
```python
def parallel_workflow(queries):
with ThreadPoolExecutor() as executor:
results = list(executor.map(
search_agent.run,
queries
))
return merge_results(results)
```
**Characteristics**:
- Fast (parallel)
- More complex
- Independent subtasks needed
- Used in 30% of workflows
---
### Pattern 3: Conditional (Branching)
```mermaid
graph TD
A["Input"] --> B{Task Type?}
B -->|Type A| C["Agent A"]
B -->|Type B| D["Agent B"]
B -->|Type C| E["Agent C"]
C --> F["Result"]
D --> F
E --> F
```
**Implementation**:
```python
def conditional_routing(task):
task_type = classifier.predict(task)
if task_type == "research":
return research_agent.run(task)
elif task_type == "analysis":
return analysis_agent.run(task)
else:
return general_agent.run(task)
```
**Characteristics**:
- Dynamic routing
- Specialist agents
- Good for diverse inputs
- Used in 40% of workflows
---
### Pattern 4: Feedback Loop
```mermaid
graph TD
A["Agent Work"] --> B["Quality Check"]
B --> C{OK?}
C -->|Yes| D["Return Result"]
C -->|No| E["Re-do Work"]
E --> A
```
**Implementation**:
```python
def feedback_loop(task):
max_iterations = 3
for i in range(max_iterations):
result = agent.run(task)
quality = evaluator.assess(result)
if quality > threshold:
return result
# Ask LLM to improve
task = f"Improve on: {result}\nFeedback: {quality.feedback}"
return result
```
**Characteristics**:
- Quality improvement
- Iterative refinement
- Expensive (multiple passes)
- Used in 25% of workflows
---
## Comparison: Which Strategy to Use?
```
Sequential Parallel Message Event Hierarchical
(State) (State) Passing Bus (Manager)
─────────────────────────────────────────────────────────────────────
Complexity Low Low Medium High Medium
Consistency Strong Strong Eventual Eventual Strong
Latency Slow Fast Medium Medium Medium
Scalability Poor Medium Good Good Good
Debugging Easy Medium Hard Hard Medium
Distributed Poor Poor Good Good Medium
Failures Sequential Stop all Isolated Cascade Manager
Best For:
Simple Simple Large Complex Enterprise
sequential parallel async workflows hierarchies
```
---
## Real-World Coordination Scenarios
### Scenario 1: Research Paper Analysis (Sequential)
```
Coordinator decides:
1. Search for papers (Search Agent)
2. Fetch full texts (Fetch Agent)
3. Summarize each (Summarize Agent)
4. Compile report (Report Agent)
Why Sequential:
• Step 2 depends on Step 1
• Natural ordering
• Easy to debug
```
### Scenario 2: Customer Support (Parallel + Routing)
```
Coordinator decides:
IF issue_type == "billing":
→ Billing Agent
ELIF issue_type == "technical":
→ Technical Agent
ELSE:
→ General Agent
Multiple requests processed in parallel
Why Parallel + Routing:
• Incoming requests independent
• Route to specialist
• Process simultaneously
```
### Scenario 3: Large Enterprise Project (Hierarchical)
```
Project Manager
- Research Team Lead
- Literature Agent
- Patent Agent
- Industry Agent
- Analysis Team Lead
- Competitive Analysis Agent
- Market Agent
- Technical Agent
- Reporting Lead
- Visualization Agent
- Documentation Agent
Why Hierarchical:
• Large, complex project
• Multiple teams
• Clear reporting structure
• Distributed decisions
```
---
## Failure Modes and Recovery
### Sequential Coordination
**Failure**: Agent A fails
**Impact**: Entire pipeline stops
**Recovery**:
```python
try:
result = search_agent.run(query)
except SearchError:
result = use_cached_results()
```
---
### Parallel Coordination
**Failure**: One parallel task fails
**Impact**: Partial results from others
**Recovery**:
```python
results = []
for future in futures:
try:
results.append(future.result())
except Exception:
results.append(None) # Partial result
```
---
### Message Passing
**Failure**: Message lost
**Impact**: Dependent agent waits
**Recovery**:
```python
try:
message = queue.receive(timeout=30)
except TimeoutError:
message = retry_send_request()
```
---
### Event-Driven
**Failure**: Event processing agent crashes
**Impact**: Event lost, downstream agents wait
**Recovery**:
```python
try:
handle_event(event)
mark_event_processed()
except Exception:
requeue_event() # Retry later
```
---
## Choosing Coordination Strategy
### Decision Framework
```
START
│
- Single agent sufficient?
- YES → No coordination needed
- NO → Continue
│
- Tasks sequential?
- YES → Use Shared State or Message Passing
- NO → Continue
│
- Tasks independent?
- YES → Use Parallel + Routing
- NO → Continue
│
- Many different agent types?
- YES → Use Hierarchical or Event Bus
- NO → Continue
│
- Distributed deployment?
YES → Use Message Passing or Event Bus
NO → Use Shared State
```
### Quick Selection
| Scenario| Strategy| Why|
|----------|----------|-----|
| Small team, sequential| Shared State| Simple, consistent|
| Parallel independent tasks| Parallel + Message| Fast, scalable|
| Many task types| Routing + Hierarchical| Clear control|
| Large distributed| Event Bus| Loose coupling|
| Real-time, many agents| Event Bus| Reactive|
| Mixed workload| Hybrid (see below)| Flexible|
---
## Hybrid Strategies (Production Norm)
Most systems combine strategies:
```
Hybrid Example 1: Sequential + Message Passing
Step 1: Shared state (fast, simple)
Step 2: Message passing (async, decoupled)
Step 3: Shared state (merge results)
Hybrid Example 2: Hierarchical + Event Bus
Manager decomposes (hierarchical)
Teams coordinate (event bus)
Results aggregated (shared state)
Hybrid Example 3: Parallel + Routing + Event Bus
Incoming requests routed (routing)
Different agents handle (parallel)
Results published (event bus)
Dashboard updates (event subscribers)
```
-
## Production Checklist
When implementing coordination:
```
Design:
Clear coordination model chosen
Data flow documented
Failure modes identified
Recovery strategies designed
Implementation:
Communication channels set up
Message/event formats defined
Serialization/deserialization working
Error handling in place
Testing:
Happy path works
Agent failure handled
Communication failure handled
Cascade failures tested
Monitoring:
Message/event logging
Latency tracking
Failure alerts
Dashboard for visualization
```
---
## Key Takeaways
1. **Shared state** = Simple, consistent, not scalable
2. **Message passing** = Scalable, async, eventual consistency
3. **Event-driven** = Loose coupling, reactive, complex
4. **Hierarchical** = Clear structure, bottleneck risk
5. **Most systems hybrid** = Combine strategies
6. **Failure handling critical** = Plan for breakdowns
7. **Observability essential** = Log all coordination
-
## Next Steps
- [Read Graph Based Orchestration](/01-agent-design/03-architecture/04-graph-based-orchestration/) - DAG workflows
- [Read Hierarchical Systems](/01-agent-design/03-architecture/05-hierarchical-agents/) - Manager-worker patterns
-
**Last Updated**: August 9, 2026