Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.
Rule of Recognition
PartialClaimed · partialDefinition · H. L. A. Hart, The Concept of Law (1961)
This meta-rule states which rules count as valid law, and which rule wins when a conflict between two rules emerges. Without it, a mere list of rules does not indicate what is actually binding.
Analysis
Only Policy-as-Prompt supplies anything like a pedigree test: a node becomes an enforced guardrail because it was extracted from an approved product, design or code document, and the link back to that text is kept. That is a criterion of standing, but it is build-time provenance for an auditor rather than a stated meta-rule the system applies, and it stops short of priority — the paper itself leaves two conflicting documents yielding two live nodes with nothing saying which controls. The other three run the opposite way: Constitutional Classifiers holds an unordered set of permitted and restricted content, SafetyAnalyst replaces rules with an aggregation of 28 weights so conflicts are settled by arithmetic rather than by a ranking, and Llama Guard makes the taxonomy a run-time argument, which answers 'what is valid law here' with 'whatever this call's prompt says'.
Open Texture
AbsentNot claimed · out of scopeDefinition · H. L. A. Hart, The Concept of Law (1961)
This stated procedure governs who and how decides unclear cases at the edges of stated rules, when it’s unclear whether and how it applies to the case at hand.
Analysis
None of the four separates a clear application from a borderline one, so there is no point at which a penumbral case is recognised as such and handed to anyone. Llama Guard and the Policy-as-Prompt classifiers return a verdict with no account of it; Constitutional Classifiers resolves the edge with a threshold; SafetyAnalyst's harm-benefit tree is a written and inspectable procedure, but it is the ordinary procedure applied uniformly to every input, not a procedure for what happens when the rules run out, and it makes vagueness disappear into a score rather than assigning it to a decider. The closest designated human, Policy-as-Prompt's human-in-the-loop conformity review, sits over the policy set before deployment and never sees an individual hard case.
Principles of Legality
PartialClaimed · partialDefinition · Lon L. Fuller, The Morality of Law (1964)
This set of properties establishes what makes governing by rules legitimate: rules must be public, clear, non-contradictory, stable over time, possible to follow and, crucially, fit “congruence” (i.e., must match how they end up being enforced).
Analysis
Splitting the rule text from the enforcing organ is what makes congruence checkable at all, and two papers get real purchase on it: Policy-as-Prompt compiles each classifier from cited source text so the run-time rule is traceably the written one, and Llama Guard carries its taxonomy in the prompt at decision time with weights and taxonomy released openly. What is missing is everything downstream of that: Constitutional Classifiers reports jailbreak resistance and a 0.38% refusal delta rather than any measured fit between a provision and what got blocked, Policy-as-Prompt's rules live in internal company documents the governed never see, no paper checks its rule set for contradiction, Llama Guard's swappable-per-call taxonomy is stability inverted, and SafetyAnalyst's 28 numbers give a blocked person nothing to have followed or to be shown to have breached.
Service Conception of Authority
AbsentNot claimed · out of scopeDefinition · Joseph Raz, The Morality of Freedom (1986)
This concept states that an authority’s rules deserve deference only if following them helps the governed act on reasons they already have. Authority is a service to the governed, not a power over them.
Analysis
In all four the standard is set by whoever runs the guard — an organisation's own PRDs, a lab's constitution, a deployer-supplied taxonomy, a weight vector held by the operator — and applied to a user who is not shown it and has no way to contest it. SafetyAnalyst's tree does score benefits alongside harms and names stakeholders, but that is third-party welfare accounting inside a moderation decision, not a showing that the person blocked would better act on their own reasons by deferring; the paper invokes a community's values without naming any community or any procedure by which it would ever set the 28 weights. Nothing in the strategy asks the question Raz's test asks, and no paper's evaluation would register the answer either way.