Constitutional AI: Harmlessness from AI Feedback¶
Authors: Bai et al. Year: 2022 ArXiv/Link: https://arxiv.org/abs/2212.08073
Summary¶
Constitutional AI approach using AI feedback against a set of constitutional principles to improve safety and alignment.
Key Concepts¶
- Constitutional principles
- AI feedback
- Harmlessness
- RLHF alternative
- Safety training
Impact¶
Novel approach to AI alignment using constitutional methods
Category¶
Claude Models