Skip to content

Coordination Strategies

Overview

When you have multiple agents, they need to coordinate. How do they communicate? How do they share information? How do they make decisions together?

This file covers the proven strategies for agent coordination in production systems (2025-2026).


The Coordination Problem

Multiple Agents:
 Agent A Agent B
 ↓ ↓
 - Must coordinate ─→
- ← Share results ←─

Challenges:
 • Information sharing
 • Decision making
 • Conflict resolution
 • Fault handling
 • Consistency
```

-

## 4 Core Coordination Strategies

### Strategy 1: Shared State (Centralized)

**How it works**: All agents read/write to shared database

```python
class SharedStateCoordination:
 def __init__(self):
 self.shared_state = Database()

 def agent_a_work(self):
 # Read state
 data = self.shared_state.get("research_data")

 # Process
 processed = self.process(data)

 # Write back
 self.shared_state.put("processed_data", processed)

 def agent_b_work(self):
 # Read from agent A's output
 processed = self.shared_state.get("processed_data")

 # Use it
 analysis = self.analyze(processed)

 # Write result
 self.shared_state.put("analysis", analysis)
```

**Pros**:

- Simple to implement
- Guaranteed consistency
- Easy to debug
- Natural for sequential work

**Cons**:

- Central point of failure
- Scalability bottleneck
- Locking issues (who modifies what?)
- Not good for distributed systems

**When to use**:

- Small teams (2-5 agents)
- Sequential workflows
- Single-machine deployment
- Consistency critical

**Production systems**: 40% use this

---

### Strategy 2: Message Passing (Asynchronous)

**How it works**: Agents send messages to each other via queue

```python
class MessagePassingCoordination:
 def __init__(self):
 self.message_queue = MessageQueue()

 def agent_a_work(self):
 # Do work
 result = self.research(query)

 # Send message to Agent B
 self.message_queue.send(
 to="agent_b",
 message={"type": "research_complete", "data": result}
)

 def agent_b_work(self):
 # Wait for message
 message = self.message_queue.receive(from_agent="agent_a")

 if message.type == "research_complete":
 # Process Agent A's result
 analysis = self.analyze(message.data)
```

**Pros**:

- Decoupled (agents independent)
- Scalable (add agents easily)
- Distributed-friendly
- Asynchronous (responsive)

**Cons**:

- Eventual consistency (delays)
- Harder to debug
- Message ordering issues
- Lost message handling

**When to use**:

- Large teams (5-20 agents)
- Parallel work
- Distributed systems
- Real-time requirements

**Production systems**: 35% use this

-

### Strategy 3: Event-Driven (Pub-Sub)

**How it works**: Agents publish events, others subscribe

```python
class EventDrivenCoordination:
 def __init__(self):
 self.event_bus = EventBus()

 def agent_a_work(self):
 # Do work
 research_results = self.research()

 # Publish event
 self.event_bus.publish(
 event_type="research_complete",
 data=research_results
)

 def agent_b_setup(self):
 # Subscribe to events
 self.event_bus.subscribe(
 event_type="research_complete",
 handler=self.on_research_complete
)

 def on_research_complete(self, event):
 # Handle event
 analysis = self.analyze(event.data)

 # Publish new event
 self.event_bus.publish(
 event_type="analysis_complete",
 data=analysis
)
```

**Pros**:

- Loosely coupled (many-to-many)
- Scalable (add subscribers easily)
- Reactive (respond to events)
- Good for complex workflows

**Cons**:

- Ordering/causality issues
- Debugging complex flows
- Resource overhead
- Cascading failures

**When to use**:

- Complex workflows
- Many agents (10-20+)
- Real-time systems
- Need loose coupling

**Production systems**: 20% use this

---

### Strategy 4: Hierarchical (Manager-Worker)

**How it works**: One manager orchestrates workers

```python
class HierarchicalCoordination:
 def __init__(self):
 self.manager = ManagerAgent()
 self.workers = {
 "researcher": ResearcherAgent(),
 "analyst": AnalystAgent(),
 "writer": WriterAgent()
 }

 def run(self, goal):
 # Manager decomposes goal
 tasks = self.manager.decompose(goal)
 # {
 # "research": "Find papers on topic X",
 # "analysis": "Analyze findings",
 # "writing": "Write report"
 # }

 results = {}

 # Manager assigns to workers
 for task_name, task_desc in tasks.items():
 worker = self.workers[task_name]
 results[task_name] = worker.execute(task_desc)

 # Manager aggregates
 final = self.manager.aggregate(results)
 return final
```

**Pros**:

- Clear control flow
- Scalable vertically (add layers)
- Easy to understand
- Good for large projects

**Cons**:

- Manager becomes bottleneck
- Manager complexity grows
- Not good for peer collaboration
- Single point of failure (manager)

**When to use**:

- Large, complex projects
- Clear task hierarchy
- Supervised workers
- Enterprise systems

**Production systems**: 30% use this

---

## Coordination Patterns

### Pattern 1: Sequential (Pipeline)

```
Agent A → Agent B → Agent C → Result

Agent A: Search for papers
Agent B: Fetch full texts
Agent C: Analyze and summarize
```

**Implementation**:

```python
def sequential_pipeline(goal):
 result = goal
 for agent in [search_agent, fetch_agent, analyze_agent]:
 result = agent.run(result)
 return result
```

**Characteristics**:

- Simple, predictable
- Slow (sequential)
- Good error handling (stop at failure)
- Used in 50% of workflows

---

### Pattern 2: Parallel (Map-Reduce)

```mermaid
graph TD
 A["Input"] --> B["Agent A"]
 A --> C["Agent B"]
 A --> D["Agent C"]
 B --> E["Merge Results"]
 C --> E
 D --> E
```

**Implementation**:

```python
def parallel_workflow(queries):
 with ThreadPoolExecutor() as executor:
 results = list(executor.map(
 search_agent.run,
 queries
))
 return merge_results(results)
```

**Characteristics**:

- Fast (parallel)
- More complex
- Independent subtasks needed
- Used in 30% of workflows

---

### Pattern 3: Conditional (Branching)

```mermaid
graph TD
 A["Input"] --> B{Task Type?}
 B -->|Type A| C["Agent A"]
 B -->|Type B| D["Agent B"]
 B -->|Type C| E["Agent C"]
 C --> F["Result"]
 D --> F
 E --> F
```

**Implementation**:

```python
def conditional_routing(task):
 task_type = classifier.predict(task)

 if task_type == "research":
 return research_agent.run(task)
 elif task_type == "analysis":
 return analysis_agent.run(task)
 else:
 return general_agent.run(task)
```

**Characteristics**:

- Dynamic routing
- Specialist agents
- Good for diverse inputs
- Used in 40% of workflows

---

### Pattern 4: Feedback Loop

```mermaid
graph TD
 A["Agent Work"] --> B["Quality Check"]
 B --> C{OK?}
 C -->|Yes| D["Return Result"]
 C -->|No| E["Re-do Work"]
 E --> A
```

**Implementation**:

```python
def feedback_loop(task):
 max_iterations = 3
 for i in range(max_iterations):
 result = agent.run(task)

 quality = evaluator.assess(result)
 if quality > threshold:
 return result

 # Ask LLM to improve
 task = f"Improve on: {result}\nFeedback: {quality.feedback}"

 return result
```

**Characteristics**:

- Quality improvement
- Iterative refinement
- Expensive (multiple passes)
- Used in 25% of workflows

---

## Comparison: Which Strategy to Use?

```
 Sequential Parallel Message Event Hierarchical
 (State) (State) Passing Bus (Manager)
─────────────────────────────────────────────────────────────────────
Complexity Low Low Medium High Medium
Consistency Strong Strong Eventual Eventual Strong
Latency Slow Fast Medium Medium Medium
Scalability Poor Medium Good Good Good
Debugging Easy Medium Hard Hard Medium
Distributed Poor Poor Good Good Medium
Failures Sequential Stop all Isolated Cascade Manager

Best For:
 Simple Simple Large Complex Enterprise
 sequential parallel async workflows hierarchies
```

---

## Real-World Coordination Scenarios

### Scenario 1: Research Paper Analysis (Sequential)

```
Coordinator decides:
 1. Search for papers (Search Agent)
 2. Fetch full texts (Fetch Agent)
 3. Summarize each (Summarize Agent)
 4. Compile report (Report Agent)

Why Sequential:
 • Step 2 depends on Step 1
 • Natural ordering
 • Easy to debug
```

### Scenario 2: Customer Support (Parallel + Routing)

```
Coordinator decides:
 IF issue_type == "billing":
 → Billing Agent
 ELIF issue_type == "technical":
 → Technical Agent
 ELSE:
 → General Agent

Multiple requests processed in parallel

Why Parallel + Routing:
 • Incoming requests independent
 • Route to specialist
 • Process simultaneously
```

### Scenario 3: Large Enterprise Project (Hierarchical)

```
Project Manager
 - Research Team Lead
 - Literature Agent
 - Patent Agent
 - Industry Agent
 - Analysis Team Lead
 - Competitive Analysis Agent
 - Market Agent
 - Technical Agent
 - Reporting Lead
 - Visualization Agent
 - Documentation Agent

Why Hierarchical:
 • Large, complex project
 • Multiple teams
 • Clear reporting structure
 • Distributed decisions
```

---

## Failure Modes and Recovery

### Sequential Coordination

**Failure**: Agent A fails
**Impact**: Entire pipeline stops
**Recovery**:

```python
try:
 result = search_agent.run(query)
except SearchError:
 result = use_cached_results()
```

---

### Parallel Coordination

**Failure**: One parallel task fails
**Impact**: Partial results from others
**Recovery**:

```python
results = []
for future in futures:
 try:
 results.append(future.result())
 except Exception:
 results.append(None) # Partial result
```

---

### Message Passing

**Failure**: Message lost
**Impact**: Dependent agent waits
**Recovery**:

```python
try:
 message = queue.receive(timeout=30)
except TimeoutError:
 message = retry_send_request()
```

---

### Event-Driven

**Failure**: Event processing agent crashes
**Impact**: Event lost, downstream agents wait
**Recovery**:

```python
try:
 handle_event(event)
 mark_event_processed()
except Exception:
 requeue_event() # Retry later
```

---

## Choosing Coordination Strategy

### Decision Framework

```
START
 │
 - Single agent sufficient?
 - YES → No coordination needed
 - NO → Continue
 │
 - Tasks sequential?
 - YES → Use Shared State or Message Passing
 - NO → Continue
 │
 - Tasks independent?
 - YES → Use Parallel + Routing
 - NO → Continue
 │
 - Many different agent types?
 - YES → Use Hierarchical or Event Bus
 - NO → Continue
 │
 - Distributed deployment?
 YES → Use Message Passing or Event Bus
 NO → Use Shared State
```

### Quick Selection

| Scenario| Strategy| Why|
|----------|----------|-----|
| Small team, sequential| Shared State| Simple, consistent|
| Parallel independent tasks| Parallel + Message| Fast, scalable|
| Many task types| Routing + Hierarchical| Clear control|
| Large distributed| Event Bus| Loose coupling|
| Real-time, many agents| Event Bus| Reactive|
| Mixed workload| Hybrid (see below)| Flexible|

---

## Hybrid Strategies (Production Norm)

Most systems combine strategies:

```
Hybrid Example 1: Sequential + Message Passing
 Step 1: Shared state (fast, simple)
 Step 2: Message passing (async, decoupled)
 Step 3: Shared state (merge results)

Hybrid Example 2: Hierarchical + Event Bus
 Manager decomposes (hierarchical)
 Teams coordinate (event bus)
 Results aggregated (shared state)

Hybrid Example 3: Parallel + Routing + Event Bus
 Incoming requests routed (routing)
 Different agents handle (parallel)
 Results published (event bus)
 Dashboard updates (event subscribers)
```

-

## Production Checklist

When implementing coordination:

```
Design:
 Clear coordination model chosen
 Data flow documented
 Failure modes identified
 Recovery strategies designed

Implementation:
 Communication channels set up
 Message/event formats defined
 Serialization/deserialization working
 Error handling in place

Testing:
 Happy path works
 Agent failure handled
 Communication failure handled
 Cascade failures tested

Monitoring:
 Message/event logging
 Latency tracking
 Failure alerts
 Dashboard for visualization
```

---

## Key Takeaways

1. **Shared state** = Simple, consistent, not scalable
2. **Message passing** = Scalable, async, eventual consistency
3. **Event-driven** = Loose coupling, reactive, complex
4. **Hierarchical** = Clear structure, bottleneck risk
5. **Most systems hybrid** = Combine strategies
6. **Failure handling critical** = Plan for breakdowns
7. **Observability essential** = Log all coordination

-

## Next Steps

- [Read Graph Based Orchestration](/01-agent-design/03-architecture/04-graph-based-orchestration/) - DAG workflows
- [Read Hierarchical Systems](/01-agent-design/03-architecture/05-hierarchical-agents/) - Manager-worker patterns

-

**Last Updated**: August 9, 2026