Skip to content

Constitutional AI: Harmlessness from AI Feedback

Authors: Bai et al. Year: 2022 ArXiv/Link: https://arxiv.org/abs/2212.08073

Summary

Constitutional AI approach using AI feedback against a set of constitutional principles to improve safety and alignment.

Key Concepts

  • Constitutional principles
  • AI feedback
  • Harmlessness
  • RLHF alternative
  • Safety training

Impact

Novel approach to AI alignment using constitutional methods

Category

Claude Models