Skip to content

GPTQ Quantization: Complete Technical Guide

Overview

GPTQ (GPT Quantization) is a post-training quantization technique that reduces Large Language Model weights to lower precision (typically int4 or int8) while maintaining near-original accuracy. It's one of the most practical and widely-used quantization methods for LLMs.

  • Paper: "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers"
  • Authors: Frantar et al., IST Austria (2023)
  • Key Innovation: Single-GPU quantization with minimal accuracy loss
  • Impact: 4x model compression, 2-4x inference speedup
  • Adoption: Widely used in production (vLLM, Ollama, HF Hub)

The Problem: Why Quantize LLMs?

Model Size Challenges

Model Sizes (float32 full precision):

Llama 2 7B:
  - Float32 (4 bytes/param): 7B × 4 = 28GB
  - Float16 (2 bytes/param): 7B × 2 = 14GB
  - Int8 (1 byte/param): 7B × 1 = 7GB
  - Int4 (0.5 bytes/param): 7B × 0.5 = 3.5GB

Llama 2 70B:
  - Float32: 70B × 4 = 280GB ❌ (no single GPU can fit!)
  - Float16: 70B × 2 = 140GB ❌ (only 8x A100s)
  - Int8: 70B × 1 = 70GB (2x A100 with effort)
  - Int4: 70B × 0.5 = 35GB ✓ (single A100 GPU!)

Hardware Constraints

Available GPU Memory:
  - Consumer GPU (RTX 4090): 24GB
  - Data Center GPU (A100): 40GB or 80GB
  - Multiple GPUs: Expensive and complex

Without Quantization:
  - Llama 2 7B: Needs 14-16GB GPU (close call!)
  - Llama 2 13B: Needs 26-30GB GPU (multiple GPUs)
  - Llama 2 70B: Needs 140GB+ GPU (8x GPUs!)

With GPTQ (int4):
  - Llama 2 7B: Fits in 8GB (consumer GPU!)
  - Llama 2 13B: Fits in 16GB (RTX 4090)
  - Llama 2 70B: Fits in 40GB (single A100)

Inference Speed Bottlenecks

Memory Bandwidth vs Compute Power:
GPU A100: 
  - FP32 throughput: 312 TFLOPS
  - Memory bandwidth: 2 TB/s
  - Byte ratio: 2 TB/s ÷ 312 TFLOPS = 6.4 bytes/FLOP

Loading weights:
  - FP32 model: Load 4 bytes/param
  - Time to load Llama 7B: (7B × 4 bytes) / 2TB/s = 14ms
  - Compute time for 1 token: ~3ms
  - Bottleneck: Memory, not compute!

With Int4:
  - Model size: 3.5GB
  - Load time: (7B × 0.5 bytes) / 2TB/s = 1.75ms
  - Net speedup: 4x reduction in memory bandwidth needed
  - Plus: Better cache locality, fewer memory stalls

What is Quantization?

Basic Concept

Quantization reduces precision of model weights:

Float32 (Full Precision):
  - 32 bits per value
  - Range: ±3.4e38
  - Precision: ~7 decimal digits
  - Example: 0.73825949 → 0.73825949

Float16 (Half Precision):
  - 16 bits per value (2x smaller)
  - Range: ±65504
  - Precision: ~3-4 decimal digits
  - Example: 0.73825949 → 0.7383 (rounded)

Int8 (1 byte per value):
  - 8 bits per value (4x smaller)
  - Range: -128 to 127
  - Integers only
  - Needs scaling: 0.73825949 → scale to [-128, 127]

Int4 (0.5 bytes per value):
  - 4 bits per value (8x smaller!)
  - Range: -8 to 7
  - Extreme precision loss
  - Heavy quantization needed

Quantization Equation

Quantized value = round(original_value / scale)

Example:
original_value = 0.75
scale = 0.1
quantized = round(0.75 / 0.1) = round(7.5) = 8

To recover (dequantize):
recovered_value = quantized × scale = 8 × 0.1 = 0.8
error = |0.75 - 0.8| = 0.05 (small!)

Why GPTQ Works: The Key Insight

The Observation

Not all weights are equally important!

In a 7B model:
- Some weights critically affect output
- Others have minimal impact
- We can quantize "unimportant" weights more aggressively
- We can quantize "important" weights less aggressively

GPTQ's Innovation:
Quantize one layer at a time, and track which weights
matter most using the Hessian (curvature information)

Hessian-Based Importance

The Hessian matrix H tells us:
"How much does the loss increase if we perturb this weight?"

High Hessian value = Important weight (sensitive)
Low Hessian value = Unimportant weight (robust)

GPTQ Strategy:
1. Compute Hessian for the layer
2. Quantize least important weights first
3. Adjust remaining weights to compensate
4. Move to next layer

The GPTQ Algorithm: Step-by-Step

Overview: Three Phases

Phase 1: Preparation
  - Load model
  - Compute Hessian for first layer
  - Sort weights by importance

Phase 2: Quantization (per layer)
  - For each weight:
    - Calculate optimal quantization
    - Quantize the weight
    - Adjust remaining weights to minimize error
  - Move to next layer

Phase 3: Finalization
  - Save quantized model
  - Verify accuracy on calibration set

Detailed Algorithm

GPTQ Algorithm Pseudocode:

for each layer L in model:
    # Step 1: Compute Hessian (Fisher Information)
    H = Hessian(layer_L)  # (weights × weights) matrix
    H_inv = inverse(H)     # Inverse (more efficient)

    # Step 2: Prepare quantization
    quantization_info = prepare_quantization(layer_L)

    for each weight w in layer_L (in order of importance):
        # Step 3: Find optimal quantization
        # Try all possible quantized values
        q = quantize(w)  # Round to nearest int4/int8

        # Step 4: Calculate error
        error = w - (q × scale)

        # Step 5: Optimal compensation (key insight!)
        # Update all remaining weights to minimize total error
        # Using Hessian information: H_inv × error
        remaining_weights = remaining_weights - (H_inv × error)

        # Step 6: Update Hessian
        # Remove this weight from Hessian (matrix update)
        H_inv = update_hessian(H_inv, w, error)

    save_layer(layer_L, quantized=True)

Example: Quantizing a Single Weight

Layer: Linear weight matrix
Shape: (4096, 4096)

Scenario: Quantize first weight to int4

Before:
w[0,0] = 0.7382594

Step 1: Calculate Hessian for this layer
H[0,0] = 5.2  (this weight is fairly important)

Step 2: Quantize to int4
int4 range: [-8, 7]
Scale for this layer: 0.1
Quantized: round(0.7382594 / 0.1) = round(7.38) = 7
Dequantized: 7 × 0.1 = 0.7

Step 3: Calculate error
error = 0.7382594 - 0.7 = 0.0382594

Step 4: Compensate remaining weights
For each other weight w_j:
    compensation = H_inv[0, j] × error
    w_j = w_j - compensation

This spreads the quantization error across all weights
using the Hessian as a guide!

Result:
- w[0,0] is now quantized (saves 4 bits)
- Other weights adjusted to compensate (maintain accuracy)
- Net effect: Minimal accuracy loss!

Mathematical Foundation: The Hessian

Why Hessian Matters

Loss function: L = MSE(quantized_output - original_output)

First derivative (gradient):
∇L/∂w = tells us direction to minimize loss

Second derivative (Hessian):
H = ∂²L/∂w² = tells us "curvature" or sensitivity

High Hessian value:
  - Weight is in a sharp valley
  - Small change → big loss increase
  - Important weight! Quantize carefully.

Low Hessian value:
  - Weight is on a flat region
  - Small change → small loss change
  - Unimportant weight. Can quantize aggressively.

Computing the Hessian

For a neural network layer:
Input: x (activation)
Weight: W
Output: y = W^T × x

Hessian H = E[x × x^T]
This is the outer product of activations!

Efficient computation:
1. Run inference on calibration data
2. Collect activations for each layer
3. Compute H = sum(x_i × x_i^T) for each sample
4. Average across samples

Approximation (GPTQ trick):
- Only compute diagonal of H for efficiency
- Or use block-wise computation
- Makes algorithm tractable for 70B models!

Quantization Methods: INT4 vs INT8

INT4 Quantization (4-bit)

Range: -8 to 7 (16 values)
Bits per weight: 4
Model compression: 8x (float32) or 4x (float16)

Advantages:
  - Maximum compression (3.5GB for 7B model!)
  - 2-4x inference speedup
  - Fits in consumer GPU memory
  - Very practical for deployment

Disadvantages:
  - Highest quantization error
  - More challenging to maintain accuracy
  - Requires careful calibration
  - May need more calibration data

Use case: When model size is critical
  - Consumer GPU deployment
  - Mobile inference
  - Batch serving with limited memory

INT8 Quantization (8-bit)

Range: -128 to 127 (256 values)
Bits per weight: 8
Model compression: 4x (float32) or 2x (float16)

Advantages:
  - Easier to achieve high accuracy
  - Faster quantization process
  - More robust
  - Good balance of size/speed

Disadvantages:
  - Less aggressive compression (7GB for 7B model)
  - Smaller speedup (1.5-2x)
  - Still requires careful quantization

Use case: When accuracy is more critical
  - Critical production systems
  - Higher quality requirements
  - Fine-tuned models

Mixed-Precision Quantization

Strategy: Use different precision for different layers!

Observation:
- Last layers are more sensitive
- Early layers are more robust

Approach:
  - Layers 1-16: Int4 (robust)
  - Layers 17-24: Int4 (robust)
  - Layers 25-31: Int8 (sensitive)
  - Layer 32: Int8 (very sensitive)
  - Attention/Output: Int8 (critical)

Result:
- Average precision: ~5-6 bits
- Compression: ~6x
- Accuracy: Near original
- Speedup: Better than pure int8

Example sizes:
  - Llama 2 7B mixed: 4-5GB
  - Llama 2 70B mixed: 40-50GB

GPTQ in Practice: Complete Workflow

Step 1: Prepare Calibration Data

from datasets import load_dataset

# Load calibration data (small sample)
dataset = load_dataset('wikitext', 'wikitext-2-raw-v1', split='train')

# Take first 128 samples for calibration
calibration_data = dataset['text'][:128]

# Why calibration data?
  - # ├─ Compute accurate Hessian
  - # ├─ Estimate quantization error distribution
  - # ├─ Find good scale factors
  - # └─ Only need ~100-1000 samples (fast!)

Step 2: Quantize with AutoGPTQ

from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig
from transformers import AutoTokenizer

# Configuration
quantize_config = BaseQuantizeConfig(
    bits=4,                    # INT4 quantization
    group_size=128,            # Group size for scale factors
    desc_act=False,            # Desc act (advanced, usually False)
    static_groups=False,       # Static groups
    true_sequential=True,      # Quantize sequentially (accurate but slower)
    damp_percent=0.1,          # Damping for Hessian (numerical stability)
)

# Load model
model_name = "meta-llama/Llama-2-7b-hf"
model = AutoGPTQForCausalLM.from_pretrained(
    model_name,
    quantize_config=quantize_config,
    device="cuda:0"
)

tokenizer = AutoTokenizer.from_pretrained(model_name)

# Tokenize calibration data
examples = [
    tokenizer(example, return_tensors="pt") 
    for example in calibration_data
]

# Quantize! (This will take time, ~30 min for 7B on A100)
model.quantize(examples, cache_examples_on_gpu=False)

# Save quantized model
model.save_quantized(
    "./llama-2-7b-gptq",
    use_safetensors=True
)

Step 3: Use Quantized Model for Inference

from auto_gptq import AutoGPTQForCausalLM
from transformers import AutoTokenizer

# Load quantized model (much faster than quantizing!)
model = AutoGPTQForCausalLM.from_quantized(
    "./llama-2-7b-gptq",
    device="cuda:0",
    use_safetensors=True
)

tokenizer = AutoTokenizer.from_pretrained(
    "meta-llama/Llama-2-7b-hf"
)

# Generate with quantized model
prompt = "Explain quantum computing"
inputs = tokenizer(prompt, return_tensors="pt")

outputs = model.generate(
    **inputs,
    max_new_tokens=256,
    do_sample=True,
    top_p=0.95,
    temperature=0.7
)

print(tokenizer.decode(outputs[0]))

Step 4: Integration with vLLM

from vllm import LLM, SamplingParams

# vLLM automatically detects GPTQ quantization!
llm = LLM(
    model="./llama-2-7b-gptq",
    quantization="gptq",  # Explicitly specify
    dtype="half",         # Use float16 for remaining compute
    gpu_memory_utilization=0.95
)

prompts = [
    "What is machine learning?",
    "Explain deep learning",
    "How do transformers work?"
]

sampling_params = SamplingParams(
    temperature=0.7,
    top_p=0.95,
    max_tokens=256
)

outputs = llm.generate(prompts, sampling_params)

for output in outputs:
    print(output.outputs[0].text)

Performance Metrics & Accuracy Loss

Benchmark Results: Llama 2 7B

Configuration:
  - Model: Llama 2 7B
  - Calibration: 128 samples from WikiText
  - Evaluation: Common benchmarks

                Original  INT8    INT4    INT4+Damp
────────────────────────────────────────────────────
Model Size      14GB      7GB     3.5GB   3.5GB
Inference Speed 1x        1.8x    2.4x    2.3x
Memory Used     14GB      7GB     3.8GB   3.8GB

Accuracy (ARC-Challenge, higher is better):
                Original  INT8    INT4    
────────────────────────────────────────
Score           53.8%     53.1%   52.2%
Loss            0%        -0.7%   -1.6%

Accuracy (MMLU, higher is better):
                Original  INT8    INT4    
────────────────────────────────────────
Score           45.9%     45.3%   44.7%
Loss            0%        -1.3%   -2.6%

Perplexity (WikiText2, lower is better):
                Original  INT8    INT4    
────────────────────────────────────────
Perplexity      6.02      6.11    6.28
Loss            0%        +1.5%   +4.3%

Key Insight:
- INT8: Almost no loss (~1%)
- INT4: Small loss (~2-4%)
- GPTQ is better than other quant methods!

Comparison: GPTQ vs Other Methods

Method              Compression  Speed  Accuracy  Difficulty
─────────────────────────────────────────────────────────────
Full Precision      1x           1x     100%      N/A
Naive INT4          8x           4x     60-70%    Easy (bad)
Quantization Aware  8x           4x     75-85%    Hard (medium)
GPTQ INT4           8x           4x     95-98%    Hard (good!)
GPTQ INT8           4x           2x     99%       Easy (excellent)

Key: GPTQ achieves best accuracy for given compression level!

Configuration Deep Dive

Group Size

What is group_size?

Model weights are divided into groups.
Each group gets its own scale factor.

Smaller group_size:
  - More scale factors (more memory)
  - Better accuracy (fine-grained)
  - Slower inference (more overhead)
  - Typical: 128, 256

Larger group_size:
  - Fewer scale factors (less memory)
  - Worse accuracy (coarse-grained)
  - Faster inference (less overhead)
  - Typical: 1024, -1 (entire weight matrix)

Example: Weight matrix (4096 × 4096)

group_size=128:
  - Divide into (4096/128) × (4096/128) = 32 × 32 = 1024 groups
  - Each group has 128 × 128 = 16,384 weights
  - 1024 scale factors (2KB for float32)
  - Memory overhead: minimal (~0.5%)

group_size=1024:
  - Divide into (4096/1024) × (4096/1024) = 4 × 4 = 16 groups
  - Each group has 1024 × 1024 = 1M weights
  - 16 scale factors (64 bytes)
  - Memory overhead: minimal

Recommended: 128 for best accuracy

Damping Factor

What is damp_percent?

During quantization, numerical issues can occur.
The Hessian might be singular or ill-conditioned.

Damping adds regularization:
H_damped = H + damp_percent × trace(H) × I

damp_percent = 0.01:
  - Light damping (1%)
  - Better accuracy
  - May have numerical issues

damp_percent = 0.1:
  - Moderate damping (10%)
  - Good balance
  - Recommended default

damp_percent = 1.0:
  - Heavy damping (100%)
  - Safer numerically
  - Might hurt accuracy

Recommended: 0.01 for accuracy, 0.1 for stability

True Sequential

Quantization order matters!

true_sequential=True:
  - Quantize weights in order (first to last)
  - More accurate (uses latest compensation info)
  - Slower (O(n²) complexity)
  - Can take 1-2 hours for 70B

true_sequential=False:
  - Quantize in groups/blocks
  - Faster (O(n) complexity)
  - Slightly less accurate
  - Takes ~30 min for 70B

Recommended: True for best quality, False for speed

Real-World Examples

Example 1: Quantizing Llama 2 70B

Scenario:
- Hardware: Single A100 (40GB)
- Original model size: 140GB (won't fit!)
- Goal: Fit on single GPU

Solution: GPTQ INT4 quantization

Process:
1. Load model in 8-bit (uses CPU offloading): 15 min
2. Calibrate on 128 samples: 5 min
3. Quantize with GPTQ: 60-90 min (true_sequential=True)
4. Save: 5 min

Result:
  - Model size: 35GB (fits on A100!)
  - Load time: 2 min (vs would never load)
  - Inference speed: 2.5x faster
  - Accuracy loss: ~2-3% (acceptable)

Actual code:
```python
from auto_gptq import AutoGPTQForCausalLM, BaseQuantizeConfig

# Only quantize (no loading full model in GPU)
model_name = "meta-llama/Llama-2-70b-hf"

# Config for 70B
quantize_config = BaseQuantizeConfig(
    bits=4,
    group_size=128,
    true_sequential=False,  # Use False for 70B (still ~60min)
    damp_percent=0.1,
)

# Load in 8-bit first to reduce memory
model = AutoGPTQForCausalLM.from_pretrained(
    model_name,
    quantize_config=quantize_config,
    device_map="auto",  # Auto distribute across GPUs if needed
    load_in_8bit=True,  # Use 8-bit for loading
)

# Quantize (this is the slow step)
import time
start = time.time()
model.quantize(examples, cache_examples_on_gpu=False)
print(f"Quantization took {(time.time() - start) / 60:.1f} minutes")

# Save
model.save_quantized("./llama-2-70b-gptq")

Example 2: Fine-Tuned Model Quantization

Scenario:
- Start with quantized Llama 2 7B
- Fine-tune on custom data
- Re-quantize to reduce size further

Process:
1. Load quantized model
2. Fine-tune on custom data (low-rank adapters)
3. Merge LoRA weights back
4. Re-quantize (different scales, better for data)

Code:
```python
from auto_gptq import AutoGPTQForCausalLM
from peft import get_peft_model, LoraConfig, TaskType

# Load pre-quantized model
model = AutoGPTQForCausalLM.from_quantized(
    "./llama-2-7b-gptq"
)

# Add LoRA for fine-tuning
peft_config = LoraConfig(
    task_type=TaskType.CAUSAL_LM,
    r=8,
    lora_alpha=16,
    lora_dropout=0.05,
    bias="none",
)
model = get_peft_model(model, peft_config)

# Fine-tune on your data
# ... training code ...

# Merge LoRA weights
model = model.merge_and_unload()

# Re-quantize!
# Note: This dequantizes first, then requantizes
model.quantize(new_calibration_data)
model.save_quantized("./llama-2-7b-finetuned-gptq")

Example 3: Quantization for Mobile

Scenario:
- Deploy LLM on mobile (iPhone, Android)
- Extreme size constraints (< 4GB)
- Need very fast inference

Solution: Ultra-aggressive GPTQ + group size optimization

Config:
- INT3 or INT2 (requires custom implementation)
- Or INT4 with large group_size (1024)
- Add layer pruning for additional compression
- Quantize activations too (advanced)

Alternative: Use smaller model + aggressive quant
- Llama 2 7B INT4: 3.5GB (too big)
- Phi 3B INT4: 1.5GB (fits!)
- Mistral 7B INT4: 3.5GB (alternative)

Code concept:
```python
# Ultra-aggressive quantization
config = BaseQuantizeConfig(
    bits=4,
    group_size=1024,  # Larger groups (faster)
    true_sequential=False,
    desc_act=True,  # Use activation order
)

# Or use smaller model
model = AutoGPTQForCausalLM.from_pretrained(
    "microsoft/phi-3-mini",  # Only 3.8B parameters
    quantize_config=config
)

Tools and Ecosystems

AutoGPTQ Library

# Main library for GPTQ quantization
from auto_gptq import (
    AutoGPTQForCausalLM,
    BaseQuantizeConfig,
)

# Features:
  - # ├─ Easy quantization API
  - # ├─ Support for many models
  - # ├─ Multi-GPU quantization
  - # └─ Efficient inference kernels

Hugging Face Hub Integration

Many pre-quantized models available!

Examples:
  - TheBloke/Llama-2-7B-Chat-GPTQ
  - TheBloke/Mistral-7B-v0.1-GPTQ
  - TheBloke/Llama-2-70B-Chat-GPTQ
  - ... 1000+ quantized models

Search at: huggingface.co/search/full-text?q=GPTQ

Usage:
```python
from auto_gptq import AutoGPTQForCausalLM

# Download pre-quantized model!
model = AutoGPTQForCausalLM.from_quantized(
    "TheBloke/Llama-2-7B-Chat-GPTQ",
    device="cuda:0"
)

Integration with Inference Engines

vLLM:
  - Auto-detects GPTQ quantization
  - Uses optimized kernels
  - 2-3x faster than naive inference

```python
from vllm import LLM

llm = LLM("TheBloke/Llama-2-7B-Chat-GPTQ")
# Automatic!

Ollama: - Can load GPTQ models - Simple command-line interface

Text Generation WebUI (oobabooga): - Supports GPTQ loading - GUI for easy interaction

LM Studio: - Desktop app supporting GPTQ - One-click quantized model download

---

## Challenges and Limitations

### Challenge 1: Accuracy Loss with INT4
Problem: Compressing 32 bits → 4 bits is aggressive!

Solution progression: - Naive INT4: ~70% accuracy (bad) - Quantization Aware Training: ~90% (better, slow) - GPTQ INT4: ~95-98% (good!) - GPTQ INT4 + calibration: ~98%+ (excellent)

Trade-off curve: Accuracy % │ - 100│ ●─── Full precision (baseline) - 95 │ ●─── GPTQ INT4 - ╱ (calibration matters) - 90 │ ● - ╱ Other methods (slower to train) - 85 │ ● - ╱ - 80 │● │ └───────────────────── Time to quantize (GPTQ: fast!)

### Challenge 2: Model Architecture Support
GPTQ works best on: ✓ Transformer-based LLMs ✓ Models with linear layers ✓ Standard attention

May have issues with: ❌ Mixture of Experts (needs research) ❌ Custom attention patterns ❌ Very unusual architectures

Solution: Use widely-used model families - Llama - Mistral - Falcon - Qwen - And many others

### Challenge 3: Quantization Time
Time to quantize: - 7B model on A100: 20-30 min - 13B model on A100: 40-60 min - 70B model on A100: 60-120 min - Multiple A100s: Can parallelize somewhat

Trade: One-time cost for permanent speedup - Quantize once: 2 hours - Use forever: 2-4x speedup - ROI: Positive within days/weeks

Optimization: - Use true_sequential=False for speed - Parallelize across GPUs - Use smaller calibration set

---

## GPTQ vs Other Quantization Methods
Comparison:

Method Bits Speed Acc% Training Size Practical ───────────────────────────────────────────────────────────────── Full Precision 32 1x 100% N/A 14GB ❌ Float16 16 1.5x 99% N/A 7GB ✓ INT8 (Post-train) 8 2x 98% No 7GB ✓ INT8 (QAT) 8 2x 99% Yes 7GB ✓ INT4 (Naive) 4 4x 70% No 3.5GB ❌ INT4 (QAT) 4 4x 90% Yes 3.5GB ❌ (slow) GPTQ INT4 4 3x 96% No 3.5GB ✓✓✓ Iterative Quant 4 3x 94% No 3.5GB ✓ AWQ 4 3x 97% No 3.5GB ✓

Best practical option: GPTQ INT4 - High accuracy (96%+) - Fast quantization (no training needed) - 8x compression - 3-4x speedup - Widely adopted (lots of pre-quantized models)

---

## Best Practices

### ✅ Do's

1. **Use GPTQ for production LLM deployment**
2. **Start with INT4** (most practical)
3. **Use group_size=128** for best accuracy
4. **Calibrate on representative data** (different from training!)
5. **Verify accuracy on benchmarks** before deployment
6. **Use pre-quantized models** if available (save time)
7. **Combine with KV cache quantization** for max speedup
8. **Profile performance** on target hardware

### ❌ Don'ts

1. ❌ Use INT4 without proper calibration
2. ❌ Quantize without verifying accuracy
3. ❌ Use group_size=1 (will fail)
4. ❌ Expect zero accuracy loss
5. ❌ Quantize for very small models (<3B may not help)
6. ❌ Assume quantization will work for custom architectures
7. ❌ Ignore numerical stability (damping helps!)
8. ❌ Quantize training data (use different calibration data!)

---

## Performance Estimation

### Will GPTQ Help Your Use Case?
Decision tree:

Is model > 7B? - Yes: GPTQ helps (almost always) - No: Maybe (depends on constraints)

Do you have GPU memory constraints? - Yes: GPTQ definitely helps - No: Still useful for speed

Are you doing inference (not training)? - Yes: Perfect for GPTQ - No: Use LoRA for training

Expected speedup: - 2-3x for INT4 (compute still bottleneck) - 1.5-2x for INT8 (memory bandwidth matters) - Plus: Fit in smaller GPU/CPU

Expected accuracy loss: - INT8: <1% (usually acceptable) - INT4: 2-4% (usually acceptable) - Varies by task (LLM output quality, MMLU, etc.) ```


Key Takeaways

🔑 GPTQ reduces model size by 4-8x while maintaining accuracy
2-4x inference speedup (compute+memory bound together)
💾 Makes large models fit on single consumer GPUs
Best practical quantization method for LLMs
📊 INT4: 96%+ accuracy (2-3% loss acceptable)
🚀 One-time quantization cost, permanent speedup benefit


Further Reading

  • GPTQ Paper: "GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers"
  • AutoGPTQ GitHub: github.com/PanQiWei/AutoGPTQ
  • Quantization Survey: Recent papers on LLM quantization techniques
  • Comparison: GPTQ vs AWQ vs SqueezeLLM