Axolotl¶
Quick Facts¶
| Aspect | Details |
|---|---|
| Organization | Open Source Community |
| Purpose | Flexible, config-driven training |
| License | Apache 2.0 |
| Installation | pip install axolotl |
| Best For | Flexible multi-method training |
| Interface | YAML configuration |
What It Does¶
Axolotl is a user-friendly, highly-configurable training framework supporting:
- LoRA and QLoRA
- Full fine-tuning
- Multiple training methods (SFT, ORPO, DPO)
- Distributed training
- Multiple data formats
- Environment setup automation
Installation¶
pip install axolotl
# Or from source
git clone https://github.com/OpenAccess-AI-Collective/axolotl
cd axolotl
pip install -e.
Core Features¶
1. Configuration-Based Training¶
# config.yaml
base_model: meta-llama/Llama-2-7b-hf
model_type: llama
tokenizer_type: llama
# Data
dataset: alpaca
dataset_prepared_path: null
val_set_size: 0.1
# Quantization (optional)
load_in_8bit: false
load_in_4bit: false
# LoRA Configuration
lora_model_dir: null
lora_r: 16
lora_alpha: 32
lora_dropout: 0.05
lora_target_modules:
- q_proj
- v_proj
- k_proj
- o_proj
# Training
output_dir:./qlora-output
num_epochs: 3
micro_batch_size: 4
gradient_accumulation_steps: 4
warmup_steps: 100
learning_rate: 2e-4
optimizer: paged_adamw_32bit
# Advanced
deepspeed: null
flash_attention: true
2. Running Training¶
# Train from config
axolotl train config.yaml
# Inference with trained model
axolotl inference config.yaml
# Merge LoRA adapters
axolotl merge_lora config.yaml \
--lora_model_dir./qlora-output \
--output_dir./merged-model
3. Multi-Method Support¶
# For DPO training
dataset: dpo_pairs
dataset_format: dpo
# For SFT
dataset_format: sharegpt
# For ORPO
dataset_format: conversation
Memory Profiles¶
# For consumer GPU (RTX 4090, 24GB)
load_in_4bit: true
bnb_4bit_quant_type: nf4
lora_r: 16
micro_batch_size: 4
gradient_accumulation_steps: 4
# For A100 (80GB)
load_in_8bit: false
lora_r: 32
micro_batch_size: 16
gradient_accumulation_steps: 2
Data Format Support¶
# Alpaca format
dataset_format: alpaca
# ShareGPT format (conversations)
dataset_format: sharegpt
# DPO format (preferred/rejected)
dataset_format: dpo
# Conversation format
dataset_format: conversation
Distributed Training¶
# For multi-GPU training
deepspeed: configs/deepspeed.json
# Model parallel (for very large models)
world_size: 4
rank: 0
Strengths¶
Flexibility - Multiple training methods in one framework Configuration - No code needed, YAML-driven Production-Ready - Supports complex setups Data Formats - Handles various dataset formats Active Development - Regular updates
Weaknesses¶
Speed - Slower than specialized tools like Unsloth Learning Curve - Many configuration options Memory - Similar to TRL without QLoRA benefits
Best Practices¶
- Start Simple - Use basic config, then customize
- Validate Data - Use the validation split
- Monitor Training - Watch metrics during training
- Use Flash Attention - Significantly speeds up training
Common Recipes¶
Recipe 1: Quick LoRA Fine-tuning¶
lora_r: 8
lora_alpha: 16
micro_batch_size: 8
gradient_accumulation_steps: 2
num_epochs: 1
Recipe 2: QLoRA on Consumer GPU¶
load_in_4bit: true
lora_r: 16
micro_batch_size: 4
gradient_accumulation_steps: 8
num_epochs: 3
Recipe 3: Full Training on A100¶
lora_model_dir: null # No LoRA - full training
micro_batch_size: 32
gradient_accumulation_steps: 1
learning_rate: 1e-4
When to Use Axolotl¶
| Scenario | Recommendation |
|---|---|
| Config-driven training | Best |
| Multi-method experiments | Best |
| Production training | Good |
| Speed critical | Use Unsloth |
| RLHF needed | Use TRL |
Resources¶
-
Next: Try 03 Unsloth for speed-optimized training