Part 2¶
Overview¶
Tools for optimized inference, serving models in production, and multi-provider APIs.
| Tool | Purpose | Best For | Complexity |
|---|---|---|---|
| Vllm | High-performance inference | Production at scale | Medium |
| Ollama | Local model running | Development & demos | Low |
| Tensorrt Llm | NVIDIA GPU optimization | Maximum performance | High |
| Litellm | Multi-provider unified API | Cost optimization | Low |
Quick Comparison¶
Performance (Llama 2 7B, single token)
- vLLM: 200+ tokens/sec
- Ollama: 50-100 tokens/sec
- TensorRT-LLM: 250+ tokens/sec
- LiteLLM: Depends on provider
Setup Time
- Ollama (5 minutes)
- LiteLLM (5 minutes)
- vLLM (10 minutes)
- TensorRT-LLM (1-2 hours)
Memory Usage (Llama 2 7B, fp16)
- vLLM: 14-16GB
- Ollama: 15-17GB
- TensorRT-LLM: 14GB
- LiteLLM: Varies by provider
Decision Tree¶
graph TD
A["Need inference?"] -->|Local/demo| B["Ollama"]
A -->|Production| C["vLLM"]
A -->|NVIDIA GPU| D["TensorRT-LLM"]
A -->|Multi-provider| E["LiteLLM"]
C -->|Need max performance| D
Deployment Checklist¶
For vLLM¶
- GPU with 16GB+ VRAM
- FastAPI for REST API
- Docker for containerization
- Load balancer for scaling
For Ollama¶
- Any GPU or CPU
- Docker compose for scaling
- Simple HTTP API built-in
For TensorRT-LLM¶
- NVIDIA GPU (A100/H100)
- TensorRT infrastructure
- Triton Inference Server
For LiteLLM¶
- API keys for providers
- Fallback model setup
- Rate limiting configuration
-
Next: Pick an inference tool to learn more!