Taxonomy / Constitutionalism / Training the constitution into the weights

Training the constitution into the weights

The model is shown its own answers alongside the written principles and asked to criticize and rewrite them, then trained on the improved versions; a second model reading those same principles supplies the preference labels human raters would otherwise give. The document plus a model reading it replaces almost all the human labeling, and once training is finished the text no longer has to be present when the model answers. That is also why the adherence benchmarks sit here: when the rules have dissolved into the weights there is no rule left to inspect, so the only audit available is to break the published text into atomic claims and measure how often finished models violate them.

The method, against Constitutionalism

Scroll the diagram sideways to see all of it.

Also places in

  • Character alignment — The self-critique loop compiles the document into the weights, so the constitution does statutory and constitutive work at once. What the model carries at inference is a formed character rather than a rulebook it consults, which is what this regime describes; the approach is the empirical case of it.

Counts and the concept reading below use the primary regime only.

Concept Analysis: Theoretical Foundations

Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.

Rule of Recognition

AbsentNot claimed · out of scope

Definition · H. L. A. Hart, The Concept of Law (1961)

This meta-rule states which rules count as valid law, and which rule wins when a conflict between two rules emerges. Without it, a mere list of rules does not indicate what is actually binding.

Analysis

The constitution, as a flat list of rules, does not necessarily resolve conflicts between its rules, especially the core initial values. For example, when "be helpful" collides with "be harmless", no rule in the constitution generally says which one wins. The model resolves the clash implicitly inside its weights during training, and there is no record of how it did so. A legal system without this meta-rule could not tell you what its own law is. What’s more, this behavior might be largely inconsistent.

Open Texture

AbsentNot claimed · out of scope

Definition · H. L. A. Hart, The Concept of Law (1961)

This stated procedure governs who and how decides unclear cases at the edges of stated rules, when it’s unclear whether and how it applies to the case at hand.

Analysis

There are cases where many rules in a typical constitution may be genuinely hard to apply. For example, consider the rule "Avoid harmful content" in the case of settling whether a chemistry lesson that mentions explosives counts as harmful. There is generally no stated procedure for such a borderline case; the model often decides directly during inference, without even leaving a reasoning trace for how it settled the complexity. The hard cases the tradition worries most about are often the ones handled least transparently.

Principles of Legality

PartialClaimed · partial

Definition · Lon L. Fuller, The Morality of Law (1964)

This set of properties establishes what makes governing by rules legitimate: rules must be public, clear, non-contradictory, stable over time, possible to follow and, crucially, fit “congruence” (i.e., must match how they end up being enforced).

Analysis

Part of this checklist is met: the principles are general, and when the document is published they are also public. But congruence, the requirement that the model's actual behavior match the written principles, is not ensured without implementing additional alignment mechanisms.

Service Conception of Authority

AbsentNot claimed · out of scope

Definition · Joseph Raz, The Morality of Freedom (1986)

This concept states that an authority’s rules deserve deference only if following them helps the governed act on reasons they already have. Authority is a service to the governed, not a power over them.

Analysis

The document is written inside the lab and imposed on users; nothing tests whether deferring to it actually serves users' own reasons.

Concept Analysis: Newly Introduced

Self-critique training loop

Added

The tradition doesn’t exactly prescribe how written rules become behavior. Here, the model critiques and rewrites its own outputs against the document and is trained on the rewrites, so the document is compiled into the weights. This addition is what makes the approach work at all, and it is also what makes the rules disappear from view at inference time.

One document, two jobs

Added

The same constitution document is used both to train the model and to evaluate it afterwards. This doesn’t have a precedent in constitutional practice, where constitutions are applied by institutions separate from the ones that govern it.

Papers

Measured, not built

Instruments that measure this rather than instantiating it. This map is a map of methods, so they are indexed here but counted nowhere on the grid.