vLLM¶
Overview¶
vLLM is the most popular production LLM inference framework. Key innovation: continuous batching + PagedAttention. Fast, easy to use, scales from single GPU to multi-node clusters.
- Paper: "Efficient Memory Management for Large Language Model Serving with PagedAttention" (Kwon et al., 2023)
- Performance: 10-100x faster than PyTorch native
- Adoption: Industry standard (used by many production systems)
- Ease: Simple Python API, Docker containers
- Features: Continuous batching, LoRA support, multi-GPU, distributed
-
Installation & Basic Usage¶
Quick Start¶
# Install vLLM
pip install vllm
# Run a simple server
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-7b-hf
# Now you have OpenAI-compatible API!
# Use with
Python API¶
from vllm import LLM, SamplingParams
# Load model
llm = LLM(model="meta-llama/Llama-2-7b-hf")
# Set sampling parameters
sampling_params = SamplingParams(
temperature=0.7,
top_p=0.95,
max_tokens=512,
)
# Batch generate (automatic continuous batching!)
prompts = [
"The meaning of life is",
"Python is a",
"Machine learning is",
]
outputs = llm.generate(prompts, sampling_params)
# Print results
for output in outputs:
prompt = output.prompt
generated_text = output.outputs[0].text
print(f"Prompt: {prompt}")
print(f"Generated: {generated_text}\n")
-
Key Features¶
Continuous Batching¶
Problem (traditional):
- Request 1: Process 100 tokens (5 seconds)
- Request 2 arrives (waits for request 1)
- Request 2: Process 100 tokens (5 seconds)
- Total: 10 seconds
Solution (continuous batching):
- Request 1: Step 1 (20 tokens)
- Request 2 arrives: Step 1 (20 tokens)
- Request 1: Step 2 (20 tokens)
- Request 2: Step 2 (20 tokens)
-...continue interleaving...
- Total: 5.5 seconds (nearly 2x faster!)
vLLM implementation:
```python
# Automatic! Just send requests
# vLLM batches them together
import httpx
for i in range(100): # 100 concurrent requests
# vLLM automatically batches these together
# Much faster than processing sequentially
response = httpx.post(
"http://localhost:8000/v1/completions",
json={
"model": "Llama-2-7b",
"prompt": f"Question {i}: What is AI?",
"max_tokens": 50,
}
)
### PagedAttention
vLLM's key innovation: Optimize KV cache with paging
Problem (native):
- KV cache: Contiguous GPU memory
- Fragmentation: Can't fit new sequences
- Waste: Empty space in KV buffer
- Result: Memory inefficient
Solution (PagedAttention):
- Treat KV cache like virtual memory
- Pages: Fixed-size chunks (16KB typical)
- Non-contiguous allocation
- Like OS page tables!
Result:
- 2-4x more throughput
- Same GPU memory
- No fragmentation!
### LoRA Support
Run different LoRA adapters with same base model
from vllm import LLM, SamplingParams
from vllm.lora.request import LoRARequest
# Load base model once
llm = LLM(
model="meta-llama/Llama-2-7b-hf",
enable_lora=True,
max_lora_rank=16,
)
# Use different LoRA adapters per request
sampling_params = SamplingParams(temperature=0.7)
requests = [
{
"prompt": "Classify sentiment: I love this!",
"lora": LoRARequest("sentiment", "path/to/sentiment-lora"),
},
{
"prompt": "Translate: Hello world",
"lora": LoRARequest("translation", "path/to/translate-lora"),
},
]
# Batch with different LoRAs
outputs = llm.generate_with_lora(requests, sampling_params)
Benefits:
- Single base model in memory
- Multiple LoRA adapters
- Automatic batching across adapters
- Efficient multi-task serving
-
## Deployment: OpenAI-Compatible API
### Server Mode
```bash
# Start vLLM server (OpenAI API compatible)
python -m vllm.entrypoints.openai.api_server \
--model meta-llama/Llama-2-7b-hf \
--port 8000 \
--tensor-parallel-size 1 \
--gpu-memory-utilization 0.9
# Now compatible with OpenAI client!
Use with OpenAI Client¶
from openai import OpenAI
# Point to vLLM server
client = OpenAI(
api_key="not needed",
base_url="http://localhost:8000/v1",
)
# Use exactly like OpenAI API!
response = client.completions.create(
model="meta-llama/Llama-2-7b-hf",
prompt="The meaning of life is",
max_tokens=50,
)
print(response.choices[0].text)
Advantages:
- Drop-in replacement for OpenAI API
- Existing OpenAI code works
- Easy migration
- Open source alternative to OpenAI
-
## Performance Benchmarks
### Throughput Comparison
Model: LLaMA 7B, A100 GPU, Batch 32
Framework Throughput (tokens/sec) ──────────────────────────────────────── PyTorch 100 vLLM 1200 TensorRT-LLM 1500 vLLM + Quantization 2000+
Real-world impact:
Scenario: 1M tokens/day to generate
Without vLLM (PyTorch):
- 1M tokens / 100 tok/s = 10,000 seconds
- Time: ~3 hours per GPU
- Need: 8+ GPUs (=$10K hardware)
With vLLM:
- 1M tokens / 1200 tok/s = 833 seconds
- Time: ~14 minutes per GPU
- Need: 1 GPU (=$1.5K hardware)
- Savings: 85% infrastructure cost!
### Latency (First Token)
Measurement: Time to first token (empty queue)
Framework Latency Reason ───────────────────────────────────── PyTorch 200-500ms Full model forward pass vLLM 50-100ms Optimized kernels TensorRT 30-50ms Compiled kernels
Interactive impact:
Latency User Experience ────────────────────────── <100ms Feels instant 100-500ms Acceptable 500ms-1s Noticeable delay
1s Frustrating
vLLM achieves <100ms: Good for interactive!
-
## Advanced Features
### Distributed Inference
```python
from vllm import LLM
# Single GPU
llm = LLM(model="meta-llama/Llama-2-70b-hf")
# Multiple GPUs (Tensor Parallelism)
llm = LLM(
model="meta-llama/Llama-2-70b-hf",
tensor_parallel_size=4, # Split across 4 GPUs
)
# Multi-node (Distributed)
llm = LLM(
model="meta-llama/Llama-2-70b-hf",
tensor_parallel_size=8, # 2 nodes × 4 GPUs each
)
Streaming Outputs¶
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Llama-2-7b-hf")
sampling_params = SamplingParams(
temperature=0.7,
max_tokens=100,
)
# Use generator for streaming
outputs = llm.generate(
prompts=["Tell me a story"],
sampling_params=sampling_params,
use_tqdm=False,
)
# Or stream via API
# curl http://localhost:8000/v1/completions \
# -d '{..., "stream"
Custom Sampling¶
from vllm import LLM, SamplingParams
llm = LLM(model="meta-llama/Llama-2-7b-hf")
# Complex sampling config
sampling_params = SamplingParams(
n=3, # Generate 3 outputs per prompt
best_of=5, # Sample 5, return best 3
temperature=0.7,
top_p=0.9,
top_k=50,
frequency_penalty=0.0,
presence_penalty=0.0,
max_tokens=100,
stop_token_ids=[2], # Stop at token 2
)
outputs = llm.generate(prompts, sampling_params)
# Each output has multiple choices
for output in outputs:
for choice in output.outputs:
print(choice.text)
-
Docker Deployment¶
Quick Containerization¶
FROM nvidia/cuda:12.1.1-devel-ubuntu22.04
RUN pip install vllm
# Download model at build time (optional)
RUN python -c "from vllm import LLM; \
LLM('meta-llama/Llama-2-7b-hf')"
# Run server
CMD ["python", "-m", "vllm.entrypoints.openai.api_server", \
"--model", "meta-llama/Llama-2-7b-hf", \
"--host", "0.0.0.0"]
# Build image
docker build -t vllm-llama2.
# Run container
docker run --gpus all -p 8000:8000 vllm-llama2
# Now access at localhost:8000
-
Limitations & Considerations¶
Limitations:
Requires GPU (NVIDIA optimized, AMD/Intel support growing)
Model must fit in GPU memory (use quantization/distillation if needed)
Doesn't support all model types (transforms/diffusers supported, custom models harder)
CPU inference not primary focus
Solutions:
Quantization (4-bit, 8-bit)
Distillation (smaller models)
Multi-GPU (tensor parallelism)
Use TensorRT for more control
Key Takeaways¶
vLLM: Easiest production LLM inference framework Continuous batching: 10-100x throughput improvement PagedAttention: Efficient KV cache management OpenAI compatible: Drop-in replacement 50-100ms latency: Interactive user experience
-
Related Notes in Inference Frameworks Subdirectory¶
- Inference Frameworks Fundamentals - Overview and comparison
- 03 Tensorrt Llm - For maximum latency optimization
- Deepspeed Mii - Alternative for distributed inference