Skip to content

Part 2

Overview

Tools for optimized inference, serving models in production, and multi-provider APIs.

Tool Purpose Best For Complexity
Vllm High-performance inference Production at scale Medium
Ollama Local model running Development & demos Low
Tensorrt Llm NVIDIA GPU optimization Maximum performance High
Litellm Multi-provider unified API Cost optimization Low

Quick Comparison

Performance (Llama 2 7B, single token)

  • vLLM: 200+ tokens/sec
  • Ollama: 50-100 tokens/sec
  • TensorRT-LLM: 250+ tokens/sec
  • LiteLLM: Depends on provider

Setup Time

  1. Ollama (5 minutes)
  2. LiteLLM (5 minutes)
  3. vLLM (10 minutes)
  4. TensorRT-LLM (1-2 hours)

Memory Usage (Llama 2 7B, fp16)

  • vLLM: 14-16GB
  • Ollama: 15-17GB
  • TensorRT-LLM: 14GB
  • LiteLLM: Varies by provider

Decision Tree

graph TD
 A["Need inference?"] -->|Local/demo| B["Ollama"]
 A -->|Production| C["vLLM"]
 A -->|NVIDIA GPU| D["TensorRT-LLM"]
 A -->|Multi-provider| E["LiteLLM"]
 C -->|Need max performance| D

Deployment Checklist

For vLLM

  • GPU with 16GB+ VRAM
  • FastAPI for REST API
  • Docker for containerization
  • Load balancer for scaling

For Ollama

  • Any GPU or CPU
  • Docker compose for scaling
  • Simple HTTP API built-in

For TensorRT-LLM

  • NVIDIA GPU (A100/H100)
  • TensorRT infrastructure
  • Triton Inference Server

For LiteLLM

  • API keys for providers
  • Fallback model setup
  • Rate limiting configuration

-

Next: Pick an inference tool to learn more!