How do you teach a model the difference between helpful and harmful when it has no inherent sense of either? This episode dives into Constitutional AI, Anthropic's framework for training AI systems to be both useful and safe by giving them an explicit set of principles to reason from. It's a fascinating look at how alignment research is evolving beyond simple human feedback — and what it means to give an AI something like a conscience. Links: Anthropic, "Constitutional AI: Harmlessness from AI Feedback" (2022) https://arxiv.org/abs/2212.08073 Claude's Constitution https://www.anthropic.com/constitution Anthropic, "Teaching Claude Why" (2026) https://www.anthropic.com/research/teaching-claude-why