Taxonomy / Constitutionalism / Screening inputs and outputs with a separate model

Screening inputs and outputs with a separate model

Enforcement lives in a component outside the model being governed: a second model inspects prompts on the way in and answers on the way out, and blocks whatever the standard forbids. The standard itself varies — a hand-written risk taxonomy, a statement of permitted and restricted content turned into synthetic training data for a classifier, a rule tree extracted from a company's own product and design documents with every node linked back to its source, or an explicit harm-benefit weighting rather than a list of rules at all — but in every case it is written down separately from the model it guards, so it can be rewritten or swapped without retraining anything. Where the policy text comes from is a variant inside this strategy, not a strategy of its own; the shared move is that the rule and its enforcer are detachable from the thing being ruled.

The method, against Constitutionalism

Scroll the diagram sideways to see all of it.

Concept Analysis: Theoretical Foundations

Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.

Rule of Recognition

PartialClaimed · partial

Definition · H. L. A. Hart, The Concept of Law (1961)

This meta-rule states which rules count as valid law, and which rule wins when a conflict between two rules emerges. Without it, a mere list of rules does not indicate what is actually binding.

Analysis

Only Policy-as-Prompt supplies anything like a pedigree test: a node becomes an enforced guardrail because it was extracted from an approved product, design or code document, and the link back to that text is kept. That is a criterion of standing, but it is build-time provenance for an auditor rather than a stated meta-rule the system applies, and it stops short of priority — the paper itself leaves two conflicting documents yielding two live nodes with nothing saying which controls. The other three run the opposite way: Constitutional Classifiers holds an unordered set of permitted and restricted content, SafetyAnalyst replaces rules with an aggregation of 28 weights so conflicts are settled by arithmetic rather than by a ranking, and Llama Guard makes the taxonomy a run-time argument, which answers 'what is valid law here' with 'whatever this call's prompt says'.

Open Texture

AbsentNot claimed · out of scope

Definition · H. L. A. Hart, The Concept of Law (1961)

This stated procedure governs who and how decides unclear cases at the edges of stated rules, when it’s unclear whether and how it applies to the case at hand.

Analysis

None of the four separates a clear application from a borderline one, so there is no point at which a penumbral case is recognised as such and handed to anyone. Llama Guard and the Policy-as-Prompt classifiers return a verdict with no account of it; Constitutional Classifiers resolves the edge with a threshold; SafetyAnalyst's harm-benefit tree is a written and inspectable procedure, but it is the ordinary procedure applied uniformly to every input, not a procedure for what happens when the rules run out, and it makes vagueness disappear into a score rather than assigning it to a decider. The closest designated human, Policy-as-Prompt's human-in-the-loop conformity review, sits over the policy set before deployment and never sees an individual hard case.

Principles of Legality

PartialClaimed · partial

Definition · Lon L. Fuller, The Morality of Law (1964)

This set of properties establishes what makes governing by rules legitimate: rules must be public, clear, non-contradictory, stable over time, possible to follow and, crucially, fit “congruence” (i.e., must match how they end up being enforced).

Analysis

Splitting the rule text from the enforcing organ is what makes congruence checkable at all, and two papers get real purchase on it: Policy-as-Prompt compiles each classifier from cited source text so the run-time rule is traceably the written one, and Llama Guard carries its taxonomy in the prompt at decision time with weights and taxonomy released openly. What is missing is everything downstream of that: Constitutional Classifiers reports jailbreak resistance and a 0.38% refusal delta rather than any measured fit between a provision and what got blocked, Policy-as-Prompt's rules live in internal company documents the governed never see, no paper checks its rule set for contradiction, Llama Guard's swappable-per-call taxonomy is stability inverted, and SafetyAnalyst's 28 numbers give a blocked person nothing to have followed or to be shown to have breached.

Service Conception of Authority

AbsentNot claimed · out of scope

Definition · Joseph Raz, The Morality of Freedom (1986)

This concept states that an authority’s rules deserve deference only if following them helps the governed act on reasons they already have. Authority is a service to the governed, not a power over them.

Analysis

In all four the standard is set by whoever runs the guard — an organisation's own PRDs, a lab's constitution, a deployer-supplied taxonomy, a weight vector held by the operator — and applied to a user who is not shown it and has no way to contest it. SafetyAnalyst's tree does score benefits alongside harms and names stakeholders, but that is third-party welfare accounting inside a moderation decision, not a showing that the person blocked would better act on their own reasons by deferring; the paper invokes a community's values without naming any community or any procedure by which it would ever set the 28 weights. Nothing in the strategy asks the question Raz's test asks, and no paper's evaluation would register the answer either way.

Concept Analysis: Newly Introduced

Provenance as a validity criterion

Added

Every rule the guardrails enforce carries a link back to the sentence in a product document it was compiled from. Legal systems establish validity by enactment; this establishes it by traceability, which is weaker — a document can be authoritative here purely because someone wrote it and the compiler read it.

Screening every act in advance

Added

Legal orders enforce selectively and after the fact, and the discretion not to prosecute is part of how they work; prior restraint is the exception a legal system has to justify. A classifier guard inspects every prompt and every response before either reaches anyone, and the paper prices that at a 23.7% inference overhead and 0.38% more refusals in production traffic. Universal ex ante screening is not something the tradition contains, and it changes what congruence means: the gap between the rule and its enforcement, which Fuller asks us to measure, is closed by construction because nothing gets past the enforcer unreviewed.

Papers

Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents

Gauri Kholkar & Ratinder Ahuja, Sep 2025

arXiv:2509.23994MethodBuilt

Reads an organization's product requirement, design and code documents, builds a policy tree whose every node links back to the source text it came from, and compiles that tree into prompt-based classifiers that monitor an agent at run time.

Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Mrinank Sharma et al., Jan 2025

arXiv:2501.18837MethodBuilt

Trains classifier safeguards on synthetic data generated by prompting language models with natural-language rules, a constitution, specifying permitted and restricted content, and reports that in over 3,000 estimated hours of red teaming no red teamer found a universal jailbreak that could extract information from an early classifier-guarded model at a level of detail similar to an unguarded model across most target queries, at a cost of an absolute 0.38% increase in production-traffic refusals and 23.7% inference overhead.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Hakan Inan et al., Dec 2023

arXiv:2312.06674MethodBuilt

A Llama2-7b model instruction-tuned on a small hand-gathered dataset to classify both prompts and responses against a written safety risk taxonomy, released with open weights, matching or exceeding available content moderation tools on the OpenAI Moderation Evaluation dataset and ToxicChat, and able to take a different taxonomy at the input for zero-shot or few-shot use.

SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior

Jing-Jing Li et al., Oct 2024

arXiv:2410.16665MethodBuilt

Uses chain-of-thought reasoning to analyze a candidate AI behavior into a structured harm-benefit tree of the harmful and beneficial actions and effects it may lead to, each labeled for likelihood, severity and immediacy of impact on stakeholders, then aggregates them into a harmfulness score through 28 fully interpretable weight parameters, in an open-source prompt safety classifier distilled from 18.5 million harm-benefit features generated by frontier models on 19k prompts that reaches average F1 0.81 where existing moderation systems score below 0.72.