Skip to content
NakodaAI
← Research & Resources

Research paperAnthropic2022-12-15

Constitutional AI: Harmlessness from AI Feedback

By Yuntao Bai, Saurav Kadavath, Sandipan Kundu, and 35 further co-authors

Anthropic's method for training a harmless AI assistant using AI-generated feedback guided by a written set of principles (a "constitution"), rather than large volumes of human-labelled harmful examples.

Why it matters

It is one of the clearest published examples of a concrete alignment technique - not a principle or a policy, but a specific training method - and it names the exact tension the method is trying to resolve: making a model both harmless and willing to engage, rather than harmless by simply refusing everything.

Key takeaways

  • The only human oversight in the method is a written list of principles (the "constitution") the model uses to critique and revise its own outputs - not per-example human harm labels.
  • Two-phase process: a supervised phase where the model critiques and rewrites its own responses against the constitution, followed by a reinforcement learning phase using AI-generated preference judgments ("RL from AI Feedback").
  • Aims explicitly for a "non-evasive" harmless assistant - one that explains its objections to a harmful request rather than refusing it with no explanation.
  • Reports that chain-of-thought reasoning during the critique step improves both the judged quality and the transparency of the model's decisions.

Part of these reading paths

Related

Also worth reading

Policy documentUK Government (host); signed by 28 states + EU, later 29 + EU2023-11-01

The Bletchley Declaration by Countries Attending the AI Safety Summit

The declaration signed at the UK's AI Safety Summit at Bletchley Park committing signatory states - including the US, China, the EU and the UAE - to international cooperation on frontier-AI safety.

  • AI Governance & Policy
  • AI Safety & Alignment
  • UAE & MENA AI

Research paperOpenAI2023-03-15

GPT-4 Technical Report

OpenAI's technical report on GPT-4, a large multimodal model accepting image and text input, covering its benchmark performance and a substantial section on the safety evaluation and mitigation work done before release.

  • Research Foundations
  • AI Safety & Alignment