Skip to content

Vision & Multimodal: Agents with Images & Video

Overview

Modern agents aren't text-only. Claude 3.5 can see images, video, and documents.

Vision-capable agents enable entirely new use cases.


Vision Capabilities Timeline

Evolution

Claude 3 (Mar 2024): Images only
Claude 3.5 Sonnet (Jun 2024): Images + higher quality
Claude 3.5 Sonnet (Aug 2024): Video support!

Current (2025-2026): Full multimodal agents

Image Understanding in Agents

Basic Image Analysis

import base64
from anthropic import Anthropic

class VisionAgent:
    def __init__(self):
        self.client = Anthropic()

    def analyze_image(self, image_path: str, task: str):
        """Agent analyzes image"""

        # Read and encode image
        with open(image_path, "rb") as f:
            image_data = base64.standard_b64encode(f.read()).decode()

        # Send to Claude with task
        response = self.client.messages.create(
            model="claude-3-5-sonnet-20241022",
            max_tokens=1024,
            messages=[
                {
                    "role": "user",
                    "content": [
                        {
                            "type": "image",
                            "source": {
                                "type": "base64",
                                "media_type": "image/jpeg",
                                "data": image_data
                            }
                        },
                        {
                            "type": "text",
                            "text": task
                        }
                    ]
                }
            ]
        )

        return response.content[0].text

Video Understanding

Analyzing Videos

class VideoAgent:
    """Agent that watches and analyzes video"""

    def __init__(self):
        self.client = Anthropic()

    def analyze_video(self, video_path: str, task: str):
        """Claude analyzes video (new capability!)"""

        # Claude 3.5 Sonnet supports video
        # Pass URL or base64 encoded

        with open(video_path, "rb") as f:
            video_data = base64.standard_b64encode(f.read()).decode()

        response = self.client.messages.create(
            model="claude-3-5-sonnet-20241022",
            max_tokens=1024,
            messages=[
                {
                    "role": "user",
                    "content": [
                        {
                            "type": "video",
                            "source": {
                                "type": "base64",
                                "media_type": "video/mp4",
                                "data": video_data
                            }
                        },
                        {
                            "type": "text",
                            "text": task
                        }
                    ]
                }
            ]
        )

        return response.content[0].text

Vision + Tool Use

Visual Tool Calling

class VisionToolAgent:
    """Agent sees image, calls appropriate tool"""

    def __init__(self):
        self.client = Anthropic()
        self.tools = [
            {
                "name": "crop_image",
                "description": "Crop a region of image"
            },
            {
                "name": "enhance_brightness",
                "description": "Enhance image brightness"
            },
            {
                "name": "extract_text",
                "description": "Extract text from image (OCR)"
            }
        ]

    def process_image_with_tools(self, image_path: str):
        """Agent sees image, decides which tool to use"""

        with open(image_path, "rb") as f:
            image_data = base64.standard_b64encode(f.read()).decode()

        response = self.client.messages.create(
            model="claude-3-5-sonnet-20241022",
            max_tokens=1024,
            tools=self.tools,
            messages=[
                {
                    "role": "user",
                    "content": [
                        {
                            "type": "image",
                            "source": {
                                "type": "base64",
                                "media_type": "image/jpeg",
                                "data": image_data
                            }
                        },
                        {
                            "type": "text",
                            "text": "Enhance this image and extract any text"
                        }
                    ]
                }
            ]
        )

        # Claude chooses appropriate tools
        return response

Document Processing

PDF & Document Analysis

class DocumentAgent:
    """Agent analyzes documents, contracts, reports"""

    def analyze_document(self, document_path: str):
        """Agent reads and analyzes document"""

        # Documents as images/PDFs
        with open(document_path, "rb") as f:
            doc_data = base64.standard_b64encode(f.read()).decode()

        response = self.client.messages.create(
            model="claude-3-5-sonnet-20241022",
            max_tokens=4096,  # Longer for docs
            messages=[
                {
                    "role": "user",
                    "content": [
                        {
                            "type": "image",
                            "source": {
                                "type": "base64",
                                "media_type": "application/pdf",
                                "data": doc_data
                            }
                        },
                        {
                            "type": "text",
                            "text": "Summarize this document and extract key clauses"
                        }
                    ]
                }
            ]
        )

        return response.content[0].text

Vision Use Cases

Real Applications

Use Case Input Task
Quality Assurance Product photos Detect defects
Document Processing Scanned documents Extract info
Video Monitoring Security footage Detect anomalies
Medical Imaging X-rays, scans Assist diagnosis
Art Analysis Paintings Identify style/artist
Layout Review Website screenshots Review design

3 Warnings ⚠️

Warning 1: Token Cost

# ❌ WRONG
# Videos are expensive!
# 1 hour video = 100k+ tokens

# ✅ RIGHT
# Compress before sending
# Send key frames, not full video
# Or use video summary tool

Warning 2: Hallucination in Vision

# ❌ WRONG
# Trust vision analysis 100%
analysis = agent.analyze_image(image)
# But Claude might "hallucinate" details!

# ✅ RIGHT
# Verify vision results
analysis = agent.analyze_image(image)
verification = verify_against_source(analysis)

Warning 3: Privacy with Images

# ❌ WRONG
# Send sensitive images to API
personal_photo = load_image(path)
analyze(personal_photo)
# Image sent to Anthropic

# ✅ RIGHT
# Check privacy policy
# Redact sensitive info before sending
redacted = redact_pii(image)
analyze(redacted)

Last Updated: August 9, 2026