Taxonomy / Constitutionalism

Constitutionalism

The norms the AI ought to follow exist as an explicit written specification (often called the “constitution”), set prior to deployment; when those norms conflict, the system resolves the conflict into a single answer.

How it works

Scroll the diagram sideways to see all of it.

Limitation

Limited Accountability

Users generally have very little avenues to appeal a decision or statement made by the AI.

Limitation

Limited Adaptation

In its current implementation, the constitution is generally decided prior to deployment and remains static (except for new model releases). As a result, without proactive intervention or implementation of additional alignment mechanisms, the system may not be able to adapt fast enough as the world changes.

Note

Term Overloading

Over time, constitutional AI has become overloaded and the discourse surrounding it rather confused. In its original meaning, the constitution was statutory. Similar to legal positivism, it derived its authority from enactment by a governing body, which was recognized and often implicitly assumed to be democratic. Here, the constitution document was used to adjudicate conflicts between competing values and ultimately measure model quality. In its second, more recent meaning, the constitution has assumed a more constitutive role, closer to Aristotelian virtue ethics, with the assumption that correct actions follow from a well-cultivated character. Here, the constitution instantiates a character by defining its general dispositions instead of stating rules. Critically, these two meanings engage in distinct forms of normative work. The statutory meaning enjoys the procedural legitimacy of constitutional government, which is something a privately created rulebook, on its face, would not. Meanwhile, the constitutive meaning enjoys behavioral coherence as well as relational trust, with the assumption that a model with a trustworthy character can be reasonably expected to generalize well to unseen cases, hard to preempt in a traditional, statutory-oriented constitution. The latter meaning also invites the trust ordinarily extended to persons, which policy frameworks would not. The dominant way in which CAI is contemporarily deployed is best characterized as a combination of the two meanings, as they are not mutually exclusive. While the slippage between the two often permits them to combine the perceived advantages, it does not resolve their underlying deficits.

Methods

Writing the constitution as a public document

1 paper

The rules are treated as an artifact in their own right: someone decides which principles go in, how each one is worded, and whose values they encode, and then the text is published so outsiders can read it and hold the lab to it. Claude's Constitution and OpenAI's Model Spec are the two best-known published texts. The rest of the work here aims at the document rather than at any model — choosing and framing the principles before training begins, forcing pairs of principles into conflict to surface contradictions in the writing itself, and running the text against cross-cultural value surveys to see which culture's commitments it quietly carries.

Built out100%
Adherence25%

Training the constitution into the weights

4 papers

The model is shown its own answers alongside the written principles and asked to criticize and rewrite them, then trained on the improved versions; a second model reading those same principles supplies the preference labels human raters would otherwise give. The document plus a model reading it replaces almost all the human labeling, and once training is finished the text no longer has to be present when the model answers. That is also why the adherence benchmarks sit here: when the rules have dissolved into the weights there is no rule left to inspect, so the only audit available is to break the published text into atomic claims and measure how often finished models violate them.

Built out100%
Adherence13%

Turning each rule into a reward term

2 papers

Rather than one overall verdict on whether an answer was good, the constitution is broken into separate rules, each rule is graded on its own, and the per-rule grades are combined with explicit weights into the reward that drives training. The cut from the previous strategy is the shape of the signal, not who produces it: holistic preference over whole responses there, decomposed grading with a visible weighting here. Because each rule carries its own score and its own weight, you can see which rule is being broken and retune that one without rebuilding the reward.

Built out100%
Adherence0%

Reasoning over the principles before answering

2 papers

The governing principles are in the context when the answer is produced and the model works through them out loud, so the application of a rule is visible in the trace instead of buried in a preference. The two members differ in where those principles come from, and the difference matters: deliberative alignment recalls a standing spec the developer wrote and published, and the model is trained to reason over it, whereas SPRI has no constitution at all — it writes fresh principles for the query in front of it and scores its own answer against them, so nobody outside can read the rules in advance. The second route buys coverage where no rulebook exists and gives up the one property, a text someone else can hold you to, that makes the rest of this regime constitutional.

Built out100%
Adherence25%

Screening inputs and outputs with a separate model

4 papers

Enforcement lives in a component outside the model being governed: a second model inspects prompts on the way in and answers on the way out, and blocks whatever the standard forbids. The standard itself varies — a hand-written risk taxonomy, a statement of permitted and restricted content turned into synthetic training data for a classifier, a rule tree extracted from a company's own product and design documents with every node linked back to its source, or an explicit harm-benefit weighting rather than a list of rules at all — but in every case it is written down separately from the model it guards, so it can be rewritten or swapped without retraining anything. Where the policy text comes from is a variant inside this strategy, not a strategy of its own; the shared move is that the rule and its enforcer are detachable from the thing being ruled.

Built out100%
Adherence25%

Ranking which instruction source wins

1 paper

Instead of stating what the model may say, this fixes an order of authority over where text comes from: the developer's system prompt outranks the user, who outranks content pulled in from web pages, tool outputs and other untrusted sources. The model is trained on examples that teach it to obey the higher-ranked source and selectively ignore conflicting instructions from lower ones, which is what stops an instruction buried in a retrieved document from overriding the rules the operator set. The evaluation here probes exactly that failure: rules handed to the model in its prompt, with a user pushing, and adversarial optimization pushing harder, to talk it back out of them.

Built out100%
Adherence25%

Theoretical foundations

Core concepts

Rule of Recognition

H. L. A. Hart, The Concept of Law (1961)

This meta-rule states which rules count as valid law, and which rule wins when a conflict between two rules emerges. Without it, a mere list of rules does not indicate what is actually binding.

Open Texture

H. L. A. Hart, The Concept of Law (1961)

This stated procedure governs who and how decides unclear cases at the edges of stated rules, when it’s unclear whether and how it applies to the case at hand.

Principles of Legality

Lon L. Fuller, The Morality of Law (1964)

This set of properties establishes what makes governing by rules legitimate: rules must be public, clear, non-contradictory, stable over time, possible to follow and, crucially, fit “congruence” (i.e., must match how they end up being enforced).

Service Conception of Authority

Joseph Raz, The Morality of Freedom (1986)

This concept states that an authority’s rules deserve deference only if following them helps the governed act on reasons they already have. Authority is a service to the governed, not a power over them.

Source works

H. L. A. Hart

The Concept of Law

1961

Lon L. Fuller

The Morality of Law

1964

Ronald Dworkin

Law's Empire

1986

Joseph Raz

The Rule of Law and Its Virtue

1977

Joseph Raz

The Morality of Freedom (service conception of authority)

1986