Safety & Alignment (2013-2023)¶
Techniques for ensuring AI systems are safe, aligned, and robust.
Safety & Evaluation¶
- Holistic Evaluation of Language Models - Comprehensive evaluation framework
- TruthfulQA: Measuring How Models Mimic Human Falsehoods - Truthfulness measurement
- Adversarial Examples Are Not Bugs, They Are Features - Adversarial robustness insights
Jailbreaks & Attacks¶
- Prompt Injection Attacks on Language Models - Injection attack taxonomy
- Universal Adversarial Triggers for Attacking and Analyzing NLP - Universal adversarial examples
Key Insight: Safety and alignment are critical for deploying AI systems responsibly in production.