Get the 2026 ML Training Cookbook | 52 recipes — GRPO, Flow Matching, World Models, and everything in between Download Now →
Constitutional AI
Train language models to self-critique and revise their own outputs according to a written constitution, reducing harmful outputs without extensive human preference labeling. Constitutional AI replaces much of the human feedback in RLHF with a written set of principles (the “constitution”). The model first generates responses, then critiques its own outputs according to the constitution, and finally revises them. This self-supervision loop produces a dataset of (original → revised) pairs for supervised learning, followed by a standard RLHF stage using a reward model trained on constitution-grounded preferences.
Papers, code, and datasets
Want to explore this concept?
Whether you're evaluating constitutional ai for your workflow or need help implementing it, we can help.
