Skip to content

Unsloth: Ultra-Fast LoRA Training

Quick Facts

Aspect Details
Organization Open Source
Purpose 2-5x faster LoRA training
License Apache 2.0
Installation pip install unsloth[colab-new] @ git+...
Best For Speed-critical projects, budget-conscious
Hardware Any GPU, optimized for RTX 4090

What It Does

Unsloth is a drop-in replacement for Hugging Face that: - 2-5x faster LoRA training - 80% less VRAM consumption - Same quality output as standard training - Simple API - minimal code changes

Installation

# For Colab
pip install unsloth[colab-new] @ git+https://github.com/unslothai/unsloth.git

# For local setup
pip install unsloth @ git+https://github.com/unslothai/unsloth.git

Basic Usage

from unsloth import FastLanguageModel
import torch

# Load model with Unsloth
model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="unsloth/llama-2-7b-bnb-4bit",
    max_seq_length=2048,
    dtype=torch.float16,
    load_in_4bit=True,
)

# Add LoRA - 2-5x faster!
model = FastLanguageModel.get_peft_model(
    model,
    r=16,
    lora_alpha=32,
    lora_dropout=0.05,
    bias="none",
    use_gradient_checkpointing="unsloth",
    random_state=3407,
)

# Standard Hugging Face trainer (unchanged!)
from transformers import TrainingArguments, SFTTrainer

trainer = SFTTrainer(
    model=model,
    tokenizer=tokenizer,
    train_dataset=dataset,
    dataset_text_field="text",
    max_seq_length=2048,
    args=TrainingArguments(
        output_dir="./output",
        num_train_epochs=1,
        per_device_train_batch_size=4,
        gradient_accumulation_steps=4,
        warmup_steps=5,
        learning_rate=2e-4,
        fp16=True,
        logging_steps=1,
        optim="adamw_8bit",
        weight_decay=0.01,
        lr_scheduler_type="linear",
        seed=3407,
    ),
)

trainer.train()

Memory Comparison

Training Llama 2 7B:

Method VRAM Speed Quality
Standard LoRA 18GB 1x 100%
Unsloth LoRA 3.5GB 3x 100%
TRL QLoRA 4GB 1x 98%
Unsloth QLoRA 2GB 4x 98%

Speed Benchmarks

Model: Llama 2 7B, Batch Size: 4, Seq Len: 2048

Standard LoRA:
  Forward pass: 2.5s
  Backward pass: 2.5s
  Total: 5s per step

Unsloth LoRA:
  Forward pass: 0.6s
  Backward pass: 0.7s
  Total: 1.3s per step
  Speedup: 3.8x

Inference with Trained Model

# Load trained model
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
    model_name="./output",
    max_seq_length=2048,
    dtype=torch.float16,
    load_in_4bit=True,
)

# Prepare for inference (minor optimization)
FastLanguageModel.for_inference(model)

# Generate
inputs = tokenizer("What is AI?", return_tensors="pt").to(model.device)
outputs = model.generate(**inputs, max_length=100)
print(tokenizer.decode(outputs[0]))

Merging LoRA

# Merge LoRA adapters into base model
model.save_pretrained_merged(
    "merged_model",
    tokenizer,
    save_method="merged_16bit"  # or "merged_4bit"
)

# Now model is a standard Hugging Face model
# Can be used with any framework!

Supported Models

  • Llama 2 (7B, 13B, 70B)
  • Mistral 7B
  • Phi series
  • Gemma series
  • Many others (check repo)

Strengths

Speed - 2-5x faster than standard training
Memory - 80% reduction in VRAM
Simplicity - Drop-in replacement, no code changes
Quality - Exact same output quality
Cost - Train on RTX 4090 instead of A100

Weaknesses

Limited Methods - LoRA only, no full training
RLHF - Not optimized for PPO/preference training
Bleeding Edge - Actively developed, might have issues

Best Practices

  1. Use Quantization - 4-bit further reduces memory
  2. Gradient Checkpointing - Already enabled by default
  3. Merge After Training - For deployment simplicity
  4. Batch Size Tuning - Memory savings enable larger batches

Typical Training Time

Dataset: 10K examples, Llama 2 7B, 1 epoch

Hardware:
  RTX 4090 (24GB): ~2-3 hours with Unsloth
  A100 (80GB): ~1.5-2 hours standard LoRA

Result: Unsloth on RTX 4090 ≈ A100 standard training

When to Use Unsloth

Scenario Recommendation
Speed critical ✅ Best choice
Budget limited ✅ Best choice
Quick iteration ✅ Best choice
Consumer GPU ✅ Best choice
RLHF training ❌ Use TRL
Full fine-tuning ⚠️ Use Axolotl

Resources


Next: Compare with 04 Ludwig for no-code training