Taxonomy / Constitutionalism / Turning each rule into a reward term

Turning each rule into a reward term

Rather than one overall verdict on whether an answer was good, the constitution is broken into separate rules, each rule is graded on its own, and the per-rule grades are combined with explicit weights into the reward that drives training. The cut from the previous strategy is the shape of the signal, not who produces it: holistic preference over whole responses there, decomposed grading with a visible weighting here. Because each rule carries its own score and its own weight, you can see which rule is being broken and retune that one without rebuilding the reward.

The method, against Constitutionalism

Scroll the diagram sideways to see all of it.

Also places in

  • Deliberative aggregation — Between the written rules and the trained model sits a scoring procedure: a judge grades instances of behavior for compliance, and the optimizer acts on those scores rather than on the document. The norms that actually reach the weights are the ones the procedure certified.
  • Character alignment — The approach says it itself: after training, the rules are no longer explicitly kept. Codified in origin and procedural in mechanism, what it leaves at inference is a disposition to behave as the rules described, with no document left to appeal to.

Counts and the concept reading below use the primary regime only.

Concept Analysis: Theoretical Foundations

Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.

Rule of Recognition

AbsentNot claimed · out of scope

Definition · H. L. A. Hart, The Concept of Law (1961)

This meta-rule states which rules count as valid law, and which rule wins when a conflict between two rules emerges. Without it, a mere list of rules does not indicate what is actually binding.

Analysis

Rules pass through a reward function on their way into the weights. Conflicts between rules get settled by whatever the reward arithmetic happens to do during training, which no one states and no one can review.

Open Texture

AbsentNot claimed · out of scope

Definition · H. L. A. Hart, The Concept of Law (1961)

This stated procedure governs who and how decides unclear cases at the edges of stated rules, when it’s unclear whether and how it applies to the case at hand.

Analysis

Borderline cases are resolved implicitly during training. At inference there is no discretion: the answer to every hard case was frozen into the weights in advance.

Principles of Legality

AbsentClaimed · unmet

Definition · Lon L. Fuller, The Morality of Law (1964)

This set of properties establishes what makes governing by rules legitimate: rules must be public, clear, non-contradictory, stable over time, possible to follow and, crucially, fit “congruence” (i.e., must match how they end up being enforced).

Analysis

The checklist is not met: constancy over time, non-contradiction between provisions, and any audit of how well actual behavior fits the text are all missing.

Service Conception of Authority

AbsentNot claimed · out of scope

Definition · Joseph Raz, The Morality of Freedom (1986)

This concept states that an authority’s rules deserve deference only if following them helps the governed act on reasons they already have. Authority is a service to the governed, not a power over them.

Analysis

An invisible rule cannot be checked against one’s own reasons.

Concept Analysis: Newly Introduced

Rules as a reward signal

Added

Turning written rules into an automatic training-time reward function, albeit effective for optimization, removes the rules from view and appeal by users.

Papers