Law as an objective the model cannot override
0 papersThe rules the system obeys are not written by the lab but taken from a named body of actual law, and obeying them is placed above every other goal the system has, so no ordinary objective can trade the law away. The supporting work argues the legal side of that design: the law already recognizes actors that carry duties without being persons, so an agent can be made a duty-bearer and liability can be channeled back to whoever deployed it. The strategy is entirely about where the rules come from and that they win; it says nothing about what happens when two of those rules collide.
Compiling statutes into software
1 paperOfficial legal text goes in and a working piece of software comes out, sitting entirely outside the model: an act read clause by clause into measurable technical requirements with a suite that runs them, a generator that assembles a customized test environment from a shared library for the authority supervising a system, and a retrieval system over 242 regulatory documents from 68 jurisdictions that routes a question to the enacted law governing it and ranks legislation above commentary. The model's training and objectives are untouched; what changes is that written law becomes something a developer or a regulator can execute against a system or look up reliably. Every instance names one jurisdiction's law, because the move only works where the statute is specific enough to translate.
Aligning the model to local rules
4 papersThe norms and legal rules of one country, region, language, online forum, or statute are assembled into an explicit rule set, converted into training data, and used to tune either the assistant itself or a separate checker beside it, so the same input gets a different verdict depending on whose rules are in force and the system can name the rule it applied. The rule sets come from human-verified cultural norms across 50 countries, from jurisdiction-specific legal texts across 14 languages, from region-targeted generated data for Southeast Asia, and from privacy statutes turned into synthetic scenarios; PluRule tests the resulting skill head-on by asking a model which of 2,885 forum rules a comment breaks, and finds frontier models barely above a trivial baseline. What holds the cell together is that pipeline — local rule set, data generated from it, model trained on that data — not the secondary choice between tuning the assistant and bolting on a guard.
Ranking rules so the order settles conflicts
1 paperDesired behavior is written as a set of rules, each one a graded measure of how badly an outcome breaks it, and the rules are arranged in a priority order, so that when two rules pull opposite ways the ordering decides rather than a judgment made in the moment. Adapting the system to another jurisdiction means re-ordering the same rules instead of rewriting them, which is how one specification serves places that weigh the same considerations differently. The work also derives which edits to a rule set preserve the guarantees an earlier ordering already gave.
Checking an agent's authority before it acts
0 papersThe codified rule here is a permission — who this agent acts for and what that person allowed it to do — written down outside the agent in a form a counterparty can verify at the moment of the request and an investigator can follow afterward. In practice that means agent-specific credentials added to the ordinary sign-in standards, with permissions written in plain language compiled into auditable access-control settings; identifiers attached to individual running instances, so anyone dealing with one can look up what it is; and a wider program of shared protocols for attributing actions, shaping how agents interact, and providing remedy when one causes harm. Identity and delegation differ in scope, not in mechanism: one answers who this is, the other on whose authority it acts, and both are checked before the action rather than judged after it.
Making access conditional on compliance
0 papersNone of these tell a model what to do. Each finds something the governed party needs and withholds it until the rules are accepted: an international body certifies whole national jurisdictions against shared oversight standards, and certified states then bar imports whose supply chains embody AI from uncertified ones and deny them specialized hardware; governments license private regulators to compete for the business of supervising firms, so operating at all means buying oversight from one of them; and a behavioral use license attaches enumerated forbidden uses to the weights, so downloading the model means taking on the terms. What differs inside the cell is the gate and who holds it — trade and chips held by certified states, market entry held by governments through licensed regulators, the artifact itself held by whoever released it — and with it who must be persuaded and who goes to court when the rule is broken.
Keeping legally risky text out of the weights
1 paperThe legal problem is settled by how the system is built rather than by any rule it is asked to follow: the parameters are trained only on public-domain and permissively licensed text, and higher-risk material is held in a separate store consulted only when a question is answered. Because the risky text was never learned into the weights, a generation can be traced back to the sentence it came from, and a rights-holder's material can be pulled out of the store without retraining anything.