Skip to content

Inference & Deployment (2017-2024)

Techniques for efficient inference, quantization, and serving LLMs.

Quantization

Decoding Strategies

Serving & Optimization


Key Insight: Inference optimizations are critical for deploying LLMs in production with acceptable latency and cost.