Framework Comparison & Selection Guide¶
Detailed Comparison Matrix¶
Feature vLLM TensorRT DeepSpeed Ollama PyTorch
─────────────────────────────────────────────────────────────────
Setup Complexity Easy Hard Hard Easy Easy
Performance ⭐⭐⭐⭐ ⭐⭐⭐⭐⭐ ⭐⭐⭐⭐ ⭐⭐ ⭐⭐
Latency (ms) 70 40 70 150 200
Throughput (tok/s) 1200 1500 1200 300 100
Multi-GPU ✅ ✅ ✅ ❌ ⚠️
Distributed ✅ ✅ ✅ ❌ ❌
LoRA Support ✅ ❌ ❌ ❌ ✅
Quantization ✅ ✅ ✅ ✅ ✅
API Style Python Config Python CLI Python
OpenAI Compatible ✅ ❌ ❌ ❌ ❌
Docker ✅ ✅ ✅ ✅ ❌
Community Large Medium Large Growing Medium
Maturity High High High Medium High
Cost (for 100k tok) $0.30 $0.25 $0.30 $0.50 $1.00
(estimated GPU cost)
Scenario-Based Selection¶
Scenario 1: Fast Prototyping (Local Development)¶
Requirements:
- Quick setup (< 5 minutes)
- Single GPU
- Local only
- Not production
Best choice: Ollama
Why:
- One-liner: ollama run llama2
- Models auto-downloaded
- Great for testing
- Works on consumer GPU
Runner-up: PyTorch with Hugging Face
- More control
- Python API
- Familiar to ML engineers
Code:
```bash
# Ollama
ollama run llama2
# PyTorch
from transformers import pipeline
pipe = pipeline("text-generation", model="llama2")
pipe("Hello world")
### Scenario 2: Production Web Service (Single GPU)
Best choice: vLLM
Why: - Simple Python setup - OpenAI-compatible (drop-in) - Great performance - Docker-ready
Code:
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Llama-2-7b-hf")
outputs = llm.generate(prompts, SamplingParams())
# Or via OpenAI API:
# python -m vllm.entrypoints.openai.api_server
Docker:
FROM nvidia/cuda:12.1-runtime
RUN pip install vllm
CMD ["python", "-m", "vllm.entrypoints.openai.api_server", \
"--model", "meta-llama/Llama-2-7b-hf"]
### Scenario 3: Production, Multi-GPU Cluster
Best choice: vLLM with multi-GPU
Why: - Tensor parallelism built-in - LoRA support - Mature, widely tested - Great community support
Code:
from vllm import LLM, SamplingParams
# 8 GPUs
llm = LLM(
model="meta-llama/Llama-2-70b-hf",
tensor_parallel_size=8,
)
outputs = llm.generate(prompts, SamplingParams())
Performance: - Throughput: 10K+ tokens/sec - Latency: 50-100ms - Concurrent users: 100-200
### Scenario 4: Ultra-Low Latency Required (<50ms)
Best choice: TensorRT-LLM
Why: - Compiled kernels (fastest) - 30-50ms latency achievable - Proven in high-traffic systems - Best for real-time
Trade-offs: ❌ More complex setup ❌ Requires model compilation ❌ Less flexible than vLLM
Code:
# Compile model for TensorRT
trtllm-build \
--checkpoint_dir ./checkpoint/llama-7b \
--output_dir ./trt_engines/llama-7b
# Run inference
python run.py --model_path ./trt_engines/llama-7b
### Scenario 5: CPU-Only or Edge Deployment
Best choice: llama.cpp
Why: - CPU-optimized C++ - Quantization (4-bit, 8-bit) - Tiny binary (small download) - No GPU needed
Alternative: ONNX Runtime - Cross-platform support - CPU/GPU flexible - Production-ready
Code:
# llama.cpp
./main -m model.gguf -n 128 -p "Hello world"
# Or Python:
from llama_cpp import Llama
llm = Llama(model_path="./model.gguf")
output = llm("Hello world", max_tokens=128)
---
## Decision Tree: Final Choice
---
## Performance Benchmarks: Real Numbers
Results (tokens/sec):
Framework FP16 INT8 INT4 ──────────────────────────────── PyTorch 100 150 200 Ollama 150 200 250 vLLM 1200 1500 1800 TensorRT-LLM 1500 2000 2200 DeepSpeed 1000 1300 1600
Latency (first token, empty queue):
Framework Latency ────────────────────── PyTorch 200ms Ollama 150ms vLLM 75ms TensorRT-LLM 40ms DeepSpeed 80ms
Memory Usage (MB):
Framework Batch 1 Batch 32 ──────────────────────────── PyTorch 20GB 23GB vLLM 14GB 18GB TensorRT-LLM 14GB 16GB DeepSpeed 14GB 19GB
Key insight: vLLM is sweet spot (great perf, easy setup)
---
## Migration Paths
### From PyTorch to vLLM
```python
# BEFORE: PyTorch
from transformers import pipeline
pipe = pipeline("text-generation", model="llama2")
output = pipe("Hello", max_new_tokens=50)
# AFTER: vLLM (similar API)
from vllm import LLM, SamplingParams
llm = LLM(model="llama2")
output = llm.generate(["Hello"], SamplingParams(max_tokens=50))
# Minimal code changes needed!
From vLLM to TensorRT (if needed)¶
# Export vLLM model to TensorRT
# 1. Export to ONNX
# 2. Compile with TensorRT
# 3. Run with TensorRT runtime
# Tool:
trtllm-build --checkpoint_dir <vllm_model>
Cost Analysis¶
Scenario: Serve 1M tokens per day
Framework GPU Cost Time per day Cost/day
───────────────────────────────────────────────
PyTorch $1000 40 hours $1.66
Ollama $1000 10 hours $0.42
vLLM $1000 1 hour $0.04
TensorRT $1000 45 min $0.03
Annual cost:
PyTorch: $606
Ollama: $153
vLLM: $15
TensorRT: $11
Lesson: Framework choice >> Hardware choice for costs!
Key Takeaways¶
🎯 vLLM: Best all-around (performance + simplicity)
⚡ TensorRT: Best latency if setup complexity acceptable
🚀 Ollama: Best for prototyping and learning
💻 llama.cpp: Best for CPU and edge
📊 Framework choice affects cost more than hardware!
Related Notes in Inference Frameworks Subdirectory¶
- Inference Frameworks Fundamentals - Overview
- Vllm - Most popular choice
- 03 Tensorrt Llm - Best latency