Skip to content

Extended Thinking & Reasoning: o1/o1-mini

Overview

OpenAI's o1 model uses extended thinking - the LLM "thinks" before responding.

Changes cost model, latency, and agent patterns entirely.


What is Extended Thinking?

The Pattern

Traditional:
  Input → LLM → Output (instant)

Extended Thinking (o1):
  Input → LLM thinks (hidden) → Output (slower)

Why It Matters

# Regular Claude/GPT-4:
# - Reasoning visible
# - Fast (1-5 seconds)
# - Can be shallow

# o1 with Extended Thinking:
# - Reasoning hidden (internal)
# - Slower (10-60 seconds)
# - Deeper reasoning
# - Better on hard problems

Performance Improvement on Hard Tasks:
  Hard math: +20-30%
  Code generation: +15-25%
  Complex reasoning: +25-40%

Cost Model Change

Pricing Shift

Claude 3.5:
  Input: $3/1M tokens
  Output: $15/1M tokens

o1:
  Input: $15/1M tokens (5x!)
  Output: $60/1M tokens (4x!)
  BUT: Fewer tokens needed for same quality

Example:
  Claude: 1000 tokens @ $0.03 = $0.03
  o1: 200 tokens @ $0.012 = $0.0024

  o1 is 12x CHEAPER despite higher per-token cost!
  (Needs 5x fewer tokens for same answer)

o1 Agent Pattern

Limited Tool Support

class O1Agent:
    """o1 has limitations!"""

    def __init__(self):
        self.client = OpenAI()

    def run_with_limitations(self):
        """o1 currently has constraints"""

        # ❌ DOESN'T WORK:
        # - Tool use (function_calling)
        # - Vision/multimodal
        # - Streaming
        # - System prompts (limited)

        # ✅ WORKS:
        # - Text reasoning
        # - Math/code problems
        # - Complex analysis

        # Recommended: Use o1 for hard thinking,
        # then Claude for tool use

Hybrid Pattern (o1 + Claude)

Best of Both

class HybridReasoningAgent:
    """Combine o1 thinking with Claude tools"""

    def solve_complex_problem(self, problem: str):
        # Step 1: Use o1 for deep thinking
        thinking = self.think_deeply(problem)

        # Step 2: Use Claude for action
        action = self.decide_action(thinking)

        # Step 3: Execute tools with Claude
        result = self.execute_with_tools(action)

        return result

    def think_deeply(self, problem: str):
        """o1 thinks about problem"""

        response = self.o1_client.messages.create(
            model="o1",
            messages=[
                {"role": "user", "content": problem}
            ]
        )

        return response.content[0].text

    def decide_action(self, thinking: str):
        """Claude decides what to do based on thinking"""

        response = self.claude_client.messages.create(
            model="claude-3-5-sonnet",
            messages=[
                {
                    "role": "user",
                    "content": f"Based on this analysis: {thinking}\n\nWhat tool should we use?"
                }
            ]
        )

        return response.content[0].text

    def execute_with_tools(self, action: str):
        """Claude executes with tools"""

        # Tool use only works with Claude
        response = self.claude_client.messages.create(
            model="claude-3-5-sonnet",
            tools=self.tools,
            messages=[
                {"role": "user", "content": action}
            ]
        )

        return response

When to Use o1

Decision Matrix

Scenario Use o1 Use Claude
Math problems ✅ Yes ❌ No
Code generation ✅ Yes ✅ Claude faster
Tool use ❌ No ✅ Yes
Vision ❌ No ✅ Yes
Speed critical ❌ No ✅ Yes
Hard reasoning ✅ Yes ❌ No
Production agent ⚠️ Hybrid ✅ Claude + o1

3 Warnings ⚠️

Warning 1: Tool Unavailable

# ❌ WRONG
# Try to use tools with o1
response = o1.messages.create(
    model="o1",
    tools=my_tools  # Doesn't work!
)

# ✅ RIGHT
# Use o1 for thinking only
thinking = o1.messages.create(
    model="o1",
    # No tools!
)

# Then use Claude for tools
action = claude.messages.create(
    model="claude-3-5-sonnet",
    tools=my_tools,  # This works!
)

Warning 2: Latency Spike

# ❌ WRONG
# Use o1 for every request
# User waits 30+ seconds

# ✅ RIGHT
# Use o1 only when needed
if problem.difficulty > HIGH:
    thinking = o1.solve(problem)  # Take time
else:
    answer = claude.solve(problem)  # Fast

Warning 3: Cost Miscalculation

# ❌ WRONG
# Think o1 is always cheaper
# Actually uses more tokens for simple problems

# ✅ RIGHT
# Measure actual cost
cost_o1 = measure_cost(o1)
cost_claude = measure_cost(claude)

# Use cheaper option
# Usually Claude for simple, o1 for hard

Last Updated: August 9, 2026