Chain-of-Thought¶
Overview¶
Chain-of-Thought (CoT) is a prompting technique that makes the LLM's reasoning explicit by asking it to show intermediate steps.
Instead of jumping to an answer, the model reasons through the problem step-by-step, making each step visible.
The CoT Insight¶
Without Chain-of-Thought¶
Question: "A store has 3 shelves. Each shelf has 4 boxes.
Each box has 7 items. How many items total?"
LLM Output: "84 items"
[No reasoning shown]
Problem: If wrong, you don't know where it went wrong
With Chain-of-Thought¶
Question: "A store has 3 shelves. Each shelf has 4 boxes.
Each box has 7 items. How many items total?"
LLM Output:
"Let me think through this step-by-step:
1. First, I need to find how many boxes total
- 3 shelves × 4 boxes per shelf = 12 boxes total
2. Next, how many items in all boxes?
- 12 boxes × 7 items per box = 84 items total
Therefore, there are 84 items total."
Benefit:
Can verify reasoning
Can identify where error occurred
Easier to debug
Higher accuracy (often +10-30%)
Why CoT Works¶
Research shows CoT improves performance on complex tasks by:
- Decomposing problems - Breaking into manageable pieces
- Forcing careful reasoning - Can't rush to answer
- Creating checkpoints - Each step is a verification point
- Showing work - Makes reasoning verifiable and debuggable
Performance Impact:
Simple arithmetic: +5% accuracy improvement
Multi-step reasoning: +20-30% accuracy improvement
Complex logic: +40-60% accuracy improvement
Implementing Chain-of-Thought¶
Basic CoT Prompt¶
def chain_of_thought_prompt(question: str) -> str:
"""Create a CoT prompt"""
return f"""
Please solve this step-by-step, showing your work:
Question: {question}
Let me work through this:
1. First, I'll identify what we need to find...
2. Next, I'll break down the problem...
3. Then, I'll calculate...
4. Finally, I'll verify...
Answer:
"""
# Usage
question = "If I have $50 and spend $12.50 on groceries and $8.75 on gas, how much do I have left?"
prompt = chain_of_thought_prompt(question)
response = llm.generate(prompt)
# Output might be:
# "Let me work through this:
# 1. Starting amount
# 2. Spending on groceries
# 3. Spending on gas
# 4. Total spending
# 5. Remaining
Production CoT Implementation¶
class ChainOfThoughtReasoner:
def __init__(self):
self.llm = LLM()
self.logger = setup_logging()
def reason(self, question: str, max_steps: int = 10) -> dict:
"""Solve problem with explicit reasoning"""
# Step 1: Ask LLM to think step-by-step
system_prompt = """
You are a careful reasoner. For every problem:
1. Break it into clear steps
2. Show your work at each step
3. Verify intermediate results
4. State your final answer with confidence
Use this format:
Step 1: [What we know]
Step 2: [What we need to find]
Step 3-N: [Working through the problem]
Final Answer: [The conclusion]
Confidence: [0-100%]
"""
user_prompt = f"""
Problem: {question}
Think step-by-step:
"""
# Get reasoning
response = self.llm.generate(
system_prompt=system_prompt,
user_prompt=user_prompt
)
self.logger.debug(f"Raw response: {response}")
# Step 2: Parse reasoning
reasoning_steps = self.parse_steps(response)
final_answer = self.extract_answer(response)
confidence = self.extract_confidence(response)
# Step 3: Verify reasoning
verification = self.verify_reasoning(reasoning_steps)
return {
'question': question,
'reasoning_steps': reasoning_steps,
'final_answer': final_answer,
'confidence': confidence,
'verification': verification,
'is_valid': verification['is_valid']
}
def parse_steps(self, response: str) -> List[dict]:
"""Extract individual reasoning steps"""
steps = []
current_step = None
for line in response.split('\n'):
if line.startswith('Step'):
if current_step:
steps.append(current_step)
current_step = {'description': line, 'content': ''}
elif current_step:
current_step['content'] += line
if current_step:
steps.append(current_step)
return steps
def verify_reasoning(self, steps: List[dict]) -> dict:
"""Verify each step makes sense"""
verification = {
'is_valid': True,
'issues': [],
'strengths': []
}
for i, step in enumerate(steps):
# Check if step is logically sound
if not self.is_logical(step):
verification['is_valid'] = False
verification['issues'].append(
f"Step {i+1}: Logical error detected"
)
# Check if step builds on previous
if i > 0:
if not self.connects_to_previous(step, steps[i-1]):
verification['issues'].append(
f"Step {i+1}: Doesn't connect to previous step"
)
# Check if step is clear
if len(step['content'].strip()) > 200:
verification['strengths'].append(
f"Step {i+1}: Detailed explanation"
)
return verification
def is_logical(self, step: dict) -> bool:
"""Check if step is logically sound"""
# Simple check: contains logical connectors
logical_words = ['therefore', 'thus', 'because', 'since', 'so', 'resulting']
return any(word in step['content'].lower() for word in logical_words)
def connects_to_previous(self, step: dict, prev_step: dict) -> bool:
"""Check if step references previous findings"""
prev_content = prev_step['content'].lower()
step_content = step['content'].lower()
# Check for references like "from step 1" or "above"
return any(ref in step_content for ref in ['previous', 'step', 'above', 'we found'])
# Usage
reasoner = ChainOfThoughtReasoner()
result = reasoner.reason(
"If a train travels 60 mph for 2.5 hours, then 40 mph for 1.5 hours, what's the average speed?"
)
print(f"Answer: {result['final_answer']}")
print(f"Confidence: {result['confidence']}%")
print(f"Valid reasoning: {result['is_valid']}")
for i, step in enumerate(result['reasoning_steps'], 1):
print(f"\nStep {i}: {step['description']}")
print(step['content'])
-
CoT Patterns¶
Pattern 1: Simple Linear Reasoning¶
Question: "What's 15% of $200?"
Step 1: 15% means 15
Step 2: So we calculate (15/100) × $200
Step 3: = 0.15 × $200 = $30
Answer: $30
Pattern 2: Multi-Stage Reasoning¶
Question: "Design a database schema for an e-commerce site"
Step 1: Identify entities (Products, Orders, Users, etc.)
Step 2: Define relationships between entities
Step 3: Specify attributes for each entity
Step 4: Consider normalization
Step 5: Optimize for common queries
Answer: [Complete schema design]
Pattern 3: Error-Checking¶
Question: "Verify if 7 × 8 = 56"
Step 1: Let me calculate 7 × 8
Step 2: 7 × 8 = 7 groups of 8
Step 3: = 8 + 8 + 8 + 8 + 8 + 8 + 8
Step 4: = 16 + 16 + 16 + 8
Step 5: = 32 + 24
Step 6: = 56
Answer: Correct!
-
When CoT Helps Most¶
High-Benefit Cases¶
# Good for CoT:
- "Solve: 3x + 5 = 20" [Math, requires steps]
- "Analyze market trends" [Complex, multiple factors]
- "Debug this code" [Multi-step analysis]
- "Evaluate decision" [Pro/con weighing]
Low-Benefit Cases¶
# Not needed for CoT:
- "What color is the sky?" [Factual, no reasoning]
- "Summarize this text" [Simple extraction]
- "Translate to Spanish" [Direct transformation]
CoT Warnings¶
Warning 1: Hallucinated Reasoning¶
# WRONG
response = reasoner.reason(question)
if response['is_valid']:
answer = response['final_answer'] # Might be wrong anyway!
# The issue
# that's actually incorrect (hallucination)
# RIGHT
response = reasoner.reason(question)
if response['is_valid']:
# Double-check the answer
verified_answer = verify_independently(response['final_answer'])
if verified_answer != response['final_answer']:
# Hallucination detected!
flag_for_review(response)
Key Lesson: CoT improves reasoning but doesn't guarantee correctness.
Warning 2: Token Cost¶
# EXPENSIVE
question = "What's 2+2?"
response = reasoner.reason(question) # Shows 10 steps
# Token cost
# SMART
if is_complex_question(question):
response = reasoner.reason(question) # Use CoT
else:
response = llm.generate(question) # Direct answer
Key Lesson: Use CoT selectively, not for everything.
Warning 3: Length Doesn't Mean Better¶
# WRONG
long_response = reasoner.reason(
question,
force_detailed=True # Forces many steps
)
# Might add unnecessary steps that confuse more than help
# RIGHT
response = reasoner.reason(
question,
max_steps=5 # Sufficient for this problem
)
Key Lesson: Detailed doesn't mean accurate. Aim for clarity.
Warning 4: Garbage In, Garbage Out¶
# WRONG
question = "Can you generate arbitrary code?"
response = reasoner.reason(question) # Tries to reason about this
# LLM produces plausible-sounding but nonsensical reasoning
# RIGHT
question = "Write a function that reverses a string"
response = reasoner.reason(question) # Produces valid reasoning
Key Lesson: Good reasoning requires well-posed questions.
Best Practices¶
1. Use CoT Selectively¶
# Good
def answer_question(question: str) -> str:
if is_complex(question):
result = reasoner.reason(question)
return result['final_answer']
else:
return llm.generate(question) # Skip CoT for simple
# Bad
def answer_question(question: str) -> str:
result = reasoner.reason(question) # Always CoT
return result['final_answer']
2. Verify Intermediate Steps¶
# Good
response = reasoner.reason(question)
for step in response['reasoning_steps']:
if not verify_step(step):
return "Reasoning error detected"
# Bad
response = reasoner.reason(question)
return response['final_answer'] # Don't verify steps
3. Monitor Token Usage¶
# Good
response = reasoner.reason(question)
tokens_used = count_tokens(response)
if tokens_used > budget:
flag_excessive_cost(response)
# Bad
response = reasoner.reason(question)
# No tracking of cost
Real-World Example¶
# E-commerce chatbot using CoT
class ProductRecommendationAgent:
def __init__(self):
self.reasoner = ChainOfThoughtReasoner()
def recommend_product(self, customer_needs: str) -> dict:
"""Recommend product using reasoning"""
# Get reasoning about recommendation
analysis = self.reasoner.reason(f"""
Given these customer needs: {customer_needs}
What product should be recommended?
Consider:
1. What are the core requirements?
2. Which products meet these requirements?
3. What are the trade-offs?
4. Which is best overall?
""")
if analysis['is_valid']:
# Reasoning is sound
return {
'recommendation': analysis['final_answer'],
'reasoning': analysis['reasoning_steps'],
'confidence': analysis['confidence']
}
else:
# Reasoning had issues
return {
'recommendation': None,
'issues': analysis['verification']['issues'],
'escalate': True
}
# Usage
agent = ProductRecommendationAgent()
result = agent.recommend_product(
"I need a laptop for development with good keyboard and battery life"
)
Comparison: CoT vs Direct¶
| Aspect | CoT | Direct |
|---|---|---|
| Accuracy | 90-95% | 80-85% |
| Token Cost | Higher | Lower |
| Speed | Slower | Faster |
| Debuggability | Excellent | Poor |
| Best For | Complex reasoning | Simple facts |
Key Takeaways¶
- CoT makes reasoning explicit - Improves accuracy and debuggability
- Not always needed - Use selectively for complex questions
- Can hallucinate - Plausible-sounding wrong reasoning
- Token-expensive - Monitor cost and use strategically
- Requires validation - Don't trust output blindly
- Works best for multi-step - Simple questions don't benefit much
-
Next Steps¶
- Read Tree Of Thought - Exploring multiple reasoning branches
- Read Goal Oriented Reasoning - Reasoning toward objectives
-
Last Updated: August 9, 2026