Skip to content

Single Agent Architecture

Overview

Before scaling to multiple agents, you need a solid single-agent architecture. This is the foundation all other patterns build upon.

A well-designed single agent is simpler to debug, easier to test, and more predictable than complex multi-agent systems.

-

Single Agent Anatomy

graph TD
 A["INPUT INTERFACE<br/>User query, API request, task"] --> B["PERCEPTION LAYER<br/>• Parse input<br/>• Check constraints<br/>• Access memory/context"]
 B --> C["LLM REASONING LAYER<br/>• Understand goal<br/>• Decide action<br/>• Consider alternatives"]
 C --> D["PLANNING LAYER<br/>Optional<br/>• Decompose complex goals<br/>• Create execution plan<br/>• Check feasibility"]
 D --> E["TOOL INTERFACE LAYER<br/>• Select appropriate tool<br/>• Call tool with args<br/>• Handle errors"]
 E --> F["EXTERNAL SYSTEMS<br/>• APIs, databases, search, files<br/>• Real-world interaction"]
 F --> G["REFLECTION LAYER<br/>• Check if goal achieved<br/>• Evaluate result quality<br/>• Decide to loop or stop"]
 G --> H["OUTPUT INTERFACE<br/>Response to user, result, action taken"]
 G -.->|Loop if needed| C
```

---

## Design Principles for Single Agents

### 1. Modularity

Each layer is independent:

```python
class ModularAgent:
 def __init__(self):
 self.perception = PerceptionModule()
 self.reasoner = ReasoningModule()
 self.planner = PlanningModule()
 self.executor = ExecutionModule()
 self.reflector = ReflectionModule()

 def run(self, task):
 perceived = self.perception.process(task)
 decision = self.reasoner.reason(perceived)
 plan = self.planner.plan(decision)
 result = self.executor.execute(plan)
 reflection = self.reflector.reflect(result)
 return reflection.output
```

**Benefits**:

- Each layer testable independently
- Easy to debug
- Can upgrade one layer without others
- Clear responsibility

### 2. State Management

Agent maintains clear state:

```python
class AgentState:
 goal: str
 context: List[str]
 completed_steps: List[str]
 current_findings: Dict
 iteration_count: int

 def is_valid(self):
 """Check state invariants"""
 return (
 self.goal is not None and
 self.iteration_count < MAX_ITERATIONS
)
```

**Benefits**:

- Explicit (not hidden in LLM)
- Debuggable (see exact state)
- Persistable (can checkpoint)
- Testable (assert state properties)

### 3. Error Handling

Graceful degradation:

```python
def robust_execution(task):
 try:
 # Primary approach
 return execute_with_best_method(task)
 except PrimaryError:
 try:
 # Secondary approach
 return execute_with_fallback(task)
 except SecondaryError:
 # Graceful degradation
 return partial_solution(task)
```

**Benefits**:

- System doesn't crash
- Partial results better than none
- Users see what worked
- Observability of failures

### 4. Observability

Log everything:

```python
class ObservableAgent:
 def run(self, task):
 self.log_event("start", task=task)

 perceived = self.perceive(task)
 self.log_event("perceived", state=perceived)

 decision = self.reason(perceived)
 self.log_event("decision", action=decision)

 result = self.execute(decision)
 self.log_event("executed", result=result)

 reflection = self.reflect(result)
 self.log_event("reflected", output=reflection)

 return reflection
```

**Benefits**:

- Full audit trail
- Debugging support
- Performance analysis
- User transparency

---

## Common Single Agent Patterns

### Pattern 1: Simple Tool-Using Agent

```python
class SimpleToolAgent:
 def __init__(self, tools: Dict):
 self.llm = LLM()
 self.tools = tools
 self.memory = []

 def run(self, goal: str) -> str:
 context = f"Goal: {goal}\n"

 while True:
 # Decide what to do
 response = self.llm.generate(
 prompt=context,
 available_tools=list(self.tools.keys())
)

 if response.type == "final_answer":
 return response.content

 # Use tool
 tool_result = self.tools[response.tool](
 **response.args
)
 context += f"\n{response.tool} result: {tool_result}"
 self.memory.append(tool_result)
```

**Use When**: Task requires simple tool calls
**Complexity**: Low
**Reliability**: High

-

### Pattern 2: Planning-Based Agent

```python
class PlanningAgent:
 def run(self, goal: str) -> str:
 # Step 1: Create plan
 plan = self.llm.create_plan(goal)
 # plan = ["Step 1:...", "Step 2:...", "Step 3:..."]

 # Step 2: Execute each step
 results = []
 for i, step in enumerate(plan):
 self.llm_context = f"Executing step {i+1}/{len(plan)}: {step}"
 result = self.execute_step(step)
 results.append(result)

 # Step 3: Compile results
 final = self.llm.compile(goal, results)
 return final
```

**Use When**: Complex multi-step tasks
**Complexity**: Medium
**Reliability**: High

---

### Pattern 3: Reflection-Enhanced Agent

```python
class ReflectionAgent:
 def run(self, goal: str) -> str:
 # Generate
 output = self.generate(goal)

 # Reflect
 critique = self.llm.critique(output)

 if critique.quality_score < 0.8:
 # Improve
 output = self.improve(output, critique)

 return output
```

**Use When**: Quality is paramount
**Complexity**: Medium
**Reliability**: Very High

---

### Pattern 4: Stateful Agent with Memory

```python
class StatefulAgent:
 def __init__(self):
 self.memory = Memory()
 self.state = AgentState()

 def run(self, goal: str) -> str:
 self.state.goal = goal

 while self.state.iteration_count < MAX_ITERATIONS:
 # Perceive with memory
 context = self._build_context()

 # Reason
 decision = self.llm.decide(context)

 # Act
 result = self.execute(decision)

 # Update state and memory
 self.state.completed_steps.append(str(decision))
 self.memory.add(result)
 self.state.iteration_count += 1

 # Reflect
 if self.goal_achieved():
 break

 return self.state.result

 def _build_context(self):
 """Build context including memory"""
 recent = self.memory.recall(self.state.goal)
 return {
 "goal": self.state.goal,
 "completed": self.state.completed_steps,
 "memory": recent
 }
```

**Use When**: Multi-turn, learning needed
**Complexity**: High
**Reliability**: High

---

## Architecture Decision Points

### Decision 1: Loop vs Single Pass

```
Single Pass (RAG-style):
 Input → Process → Output
 Latency: <1s
 Cost: Low
 Quality: Medium
 Use: Simple Q&A

Loop-Based (Agentic):
 Input → Loop → Output
 Latency: 5-60s
 Cost: Higher
 Quality: High
 Use: Complex tasks
```

### Decision 2: Planning vs Reactive

```
Reactive:
 See situation → Act → See result → Adapt

 Agent: "I see we need data. Let me search."
 Flexible, adapts to surprises
 Harder to explain

Planning-Based:
 Analyze goal → Create plan → Execute steps

 Agent: "Goal requires 3 steps. Here's the plan:"
 Predictable, easier to follow
 Less flexible
```

### Decision 3: Tool-Rich vs Tool-Light

```
Tool-Rich:
 Can do many things
 Powerful
 More complex to manage
 Higher error surface

Tool-Light:
 Simple, focused
 Easy to debug
 Limited capabilities
 Can't do everything
```

---

## Implementation Checklist

```
Core Components:
 LLM selection (model, API, cost)
 Tool definition (what can agent do)
 Tool calling (function signature specification)
 Error handling (what if tool fails?)
 State management (what's tracked?)

Optional Enhancements:
 Memory system (remember past?)
 Reflection (self-critique?)
 Planning (decompose first?)
 Routing (classify before acting?)

Production Requirements:
 Observability (log everything)
 Error recovery (graceful degradation)
 Testing framework (how verify?)
 Monitoring (track quality)
 User interface (how humans interact?)
```

---

## Scaling from Single Agent

When does single agent stop working?

```
Single Agent Sufficient:
 < 5 tools
 < 30 second latency acceptable
 1-2 task types
 < 10,000 users
 Quality < 95%

Consider Multi-Agent:
 > 5 specialized tasks
 Need different specialists
 Parallel work beneficial
 > 10,000 concurrent users
 Quality > 95% needed
```

-

## Single Agent Best Practices

### 1. Start Simple

```
Week 1: LLM + 1-2 tools + error handling
Week 2: Add memory if needed
Week 3: Add planning or reflection if quality issues
```

### 2. Test Thoroughly

```python
def test_agent():
 assert agent.run("simple query") is not None
 assert agent.run("complex task") succeeds
 assert agent.handles_errors_gracefully()
 assert agent.respects_constraints()
 assert agent.logs_everything()
```

### 3. Monitor Continuously

```
Key Metrics:
- Success rate (% goals achieved)
- Quality score (accuracy)
- Latency (time to result)
- Cost per call
- Error rate
- User satisfaction
```

### 4. Iterate Based on Feedback

```
Feedback → Metric Analysis → Identify Issue
 ↓
- ┌──────────────────┴──────────────────┐
 ↓ ↓
 Quality Issue? Speed Issue?
 │ │
 Add Reflection/CoT Remove loop/Reflection
 Use better model Simplify approach
 Improve prompts Use RAG instead
```

-

## Example: Production-Grade Single Agent

```python
class ProductionSingleAgent:
 def __init__(self, config: AgentConfig):
 self.config = config
 self.llm = LLM(model=config.model)
 self.tools = self._init_tools(config.tools)
 self.memory = Memory(backend=config.memory_backend)
 self.logger = setup_logging()

 def run(self, task: Task) -> Result:
 """Main agent loop"""
 self.logger.info(f"Starting task: {task.id}")

 state = AgentState(goal=task.goal)

 try:
 while state.iteration_count < self.config.max_iterations:
 # Perceive
 context = self._build_context(state)
 self.logger.debug(f"Context size: {len(context)}")

 # Reason
 decision = self.llm.decide(
 context=context,
 tools=list(self.tools.keys())
)
 self.logger.info(f"Decided: {decision.action}")

 # Act
 try:
 result = self.tools[decision.action](
 **decision.args
)
 self.logger.info(f"Action succeeded: {decision.action}")
 except ToolError as e:
 self.logger.warning(f"Tool error: {e}")
 result = self._handle_tool_error(e, state)

 # Update state
 state.completed_steps.append(decision)
 state.last_result = result
 state.iteration_count += 1

 # Store in memory
 self.memory.add(result)

 # Reflect
 if self._goal_achieved(state):
 self.logger.info("Goal achieved")
 break

 return self._format_result(state)

 except Exception as e:
 self.logger.error(f"Agent failed: {e}")
 return Result(success=False, error=str(e))

 def _build_context(self, state: AgentState) -> str:
 """Build prompt context"""
 recent_memory = self.memory.recall(state.goal, k=5)
 return f"""
 Goal: {state.goal}
 Completed: {state.completed_steps}
 Memory: {recent_memory}
 Iteration: {state.iteration_count}/{self.config.max_iterations}
 """

 def _handle_tool_error(self, error, state):
 """Recover from tool errors"""
 self.logger.info(f"Recovering from: {error}")

 if error.recoverable:
 return {"fallback": "Using alternative approach"}
 else:
 state.has_error = True
 return {"error": str(error)}

 def _goal_achieved(self, state: AgentState) -> bool:
 """Check if goal is achieved"""
 if state.last_result is None:
 return False

 confidence = self.llm.assess_goal_achievement(
 goal=state.goal,
 result=state.last_result
)

 return confidence > self.config.success_threshold

 def _format_result(self, state: AgentState) -> Result:
 """Format final result"""
 return Result(
 success=not state.has_error,
 output=state.last_result,
 steps=len(state.completed_steps),
 memory=self.memory.get_summary()
)
```

---

## Key Takeaways

1. **Single agent is the foundation** - Master this before multi-agent
2. **Modularity enables testing** - Each layer independent
3. **State matters** - Make it explicit, persistent
4. **Error handling is critical** - Systems will fail
5. **Observability essential** - Log everything
6. **Start simple, add as needed** - Don't over-engineer
7. **Know when to scale** - Multi-agent for specific problems

-

## Next Steps

- [Read Multi Agent Systems](/01-agent-design/03-architecture/02-multi-agent-systems/) - When to scale
- [Read Coordination Strategies](/01-agent-design/03-architecture/03-coordination-strategies/) - How to coordinate
- [Jump To Selection Framework](/01-agent-design/02-core-design-patterns/04-pattern-selection-framework/) - How to choose

-

**Last Updated**: August 9, 2026