Taxonomy / Consociational codification / Aligning the model to local rules

Aligning the model to local rules

The norms and legal rules of one country, region, language, online forum, or statute are assembled into an explicit rule set, converted into training data, and used to tune either the assistant itself or a separate checker beside it, so the same input gets a different verdict depending on whose rules are in force and the system can name the rule it applied. The rule sets come from human-verified cultural norms across 50 countries, from jurisdiction-specific legal texts across 14 languages, from region-targeted generated data for Southeast Asia, and from privacy statutes turned into synthetic scenarios; PluRule tests the resulting skill head-on by asking a model which of 2,885 forum rules a comment breaks, and finds frontier models barely above a trivial baseline. What holds the cell together is that pipeline — local rule set, data generated from it, model trained on that data — not the secondary choice between tuning the assistant and bolting on a guard.

The method, against Consociational codification

Scroll the diagram sideways to see all of it.

Concept Analysis: Theoretical Foundations

Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.

Segmental autonomy

PresentClaimed · delivered

Definition · Arend Lijphart, Democracy in Plural Societies (1977)

In a deeply divided society, each group governs its own internal affairs — its schools, its family law — within its own sphere, instead of everyone living under one uniform rule.

Analysis

Each region's cultural norms and legal policies genuinely govern the answers given in it, across 50 countries and 493 regions, and the norms were human-verified rather than inferred. This is the only built system in the map where a user's sphere determines the rule applied to them.

Choice-of-law rule

PartialClaimed · partial

Definition · Friedrich Carl von Savigny, System of the Modern Roman Law (1849)

When several bodies of rules could govern the same case, an explicit rule decides which one actually does — and states the reason. Without such a rule, "different rules for different groups" has no answer for the cases in between.

Analysis

The system does state which norm it applied and cite it, which is half of Savigny's requirement and more than AP-09 manages. The half that is missing is the case in between: a query implicating two countries' norms has no allocation rule, because the geography is read off the query rather than reasoned about.

Mutual veto and proportionality

AbsentNot claimed · out of scope

Definition · Arend Lijphart, Democracy in Plural Societies (1977)

Decisions that affect all groups require every group's consent, and representation is proportional to each group's size — so no segment can simply be outvoted on what matters most to it.

Analysis

Norms are collected from regions and compiled into training data; no region consents to the compilation, none can veto how it is represented, and representation follows the availability of documented norms rather than population.

Concept Analysis: Newly Introduced

The choice-of-law clause stated in the answer

Added

Savigny's rule operates between courts, out of sight of the parties. Here the governing norm is named inside the model's response to the user, so the allocation is visible at the moment it is applied. Nothing in the tradition puts the conflicts rule in front of the person it governs.

Papers

SafeWorld: Geo-Diverse Safety Alignment

Da Yin et al., Dec 2024

arXiv:2412.06483MethodBuilt

Builds a 2,342-query benchmark grounded in human-verified cultural norms and legal policies from 50 countries and 493 regions or ethnic groups, then trains SafeWorldLM by direct preference optimization to answer the same question differently by context and to cite the norm or policy it is applying.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

Yunhan Zhao et al., May 2026

arXiv:2605.00689MethodBuilt

Derives risk categories and fine-grained rules from jurisdiction-specific legal texts and uses them to generate a safety benchmark covering 14 languages, then builds on that benchmark a diffusion-model guardrail in two sizes, a 1.5B one for fast safe or unsafe checks and a 7B one that assesses compliance against a policy supplied to it and explains its verdict, reported as consistently outperforming 11 guardrail baselines across six existing multilingual benchmarks and its own.

SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia

Panuthep Tasawong et al., Feb 2026

arXiv:2602.01618MethodBuilt

Generates region-specific safety data for Southeast Asia with an agentic data-generation pipeline rather than by machine-translating English datasets, and trains a family of safeguard models on it that detect regionally sensitive or harmful content better than existing safeguards while maintaining strong general safety performance.

GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory

Wei Fan et al., Jun 2024

arXiv:2406.11149MethodBuilt

A framework that generates synthetic scenarios grounded in privacy statutes, using contextual integrity as the bridge between statute and situation, so that a language model given them recognizes privacy violations in real court cases.