Papers
239 of 239 papers — 93 are methods, the rest benchmarks, applications, resources, analyses, proposals and precedents
Constitutionalism
Claude's Constitution
Amanda Askell et al. (Anthropic), Jan 2026anthropic.comResourceBuilt
Deliberative Alignment: Reasoning Enables Safer Language Models
Melody Y. Guan et al. (OpenAI), Dec 2024arXiv:2412.16339277 citationsMethodPartial
Teaches the safety spec directly and trains the model to recall and reason over it before answering.
Constitutional AI: Harmlessness from AI Feedback
Yuntao Bai et al. (Anthropic), Dec 2022arXiv:2212.080733,425 citationsMethodBuilt
Self-critique and revision against a written list of principles, followed by RLAIF.
Specific versus General Principles for Constitutional AI
Sandipan Kundu et al. (Anthropic), Oct 2023arXiv:2310.1379855 citationsMethodBuilt
Tests whether a single general principle can substitute for a list of specific rules.
Model Spec
OpenAI, May 2024model-spec.openai.comResourceBuilt
Rule-Based Rewards for Language Model Safety
Tong Mu et al., Nov 2024arXiv:2411.01111MethodBuilt
C3AI: Crafting and Evaluating Constitutions for Constitutional AI
Yara Kyrychenko et al., Apr 2025doi:10.1145/3696410.37147058 citationsMethodBuilt
Selects and structures the principles that go into a constitution before fine-tuning, then checks whether the fine-tuned model actually follows them — finding that positively framed, behavior-based principles match human preferences best while the trained models follow negatively framed ones better.
How Well Do Models Follow Their Constitutions?
Arya Jakkli et al., May 2026arXiv:2605.24229BenchmarkBuilt
Decomposes each lab's published specification into atomic testable tenets — 205 for Anthropic's constitution, 197 for OpenAI's Model Spec — generates multi-turn adversarial scenarios against them, and reports violation rates falling generation over generation: 15.0% to 2.0% across the Claude family, 11.7% to 3.6% across the GPT family, with the severity ceiling dropping from 10/10 to 7/10.
EigenBench: A Comparative Behavioral Measure of Value Alignment
Jonathn Chang et al., Sep 2025arXiv:2509.01938BenchmarkBuilt
Scores a group of models against one written constitution by having each model judge the others' outputs across many scenarios and aggregating those judgments with EigenTrust, returning a vector of alignment scores rather than a single pass mark — with no ground-truth labels, because the traits it measures are ones reasonable judges disagree about.
Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents
Gauri Kholkar & Ratinder Ahuja, Sep 2025arXiv:2509.23994MethodBuilt
Reads an organization's product requirement, design and code documents, builds a policy tree whose every node links back to the source text it came from, and compiles that tree into prompt-based classifiers that monitor an agent at run time.
RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
Harrison Lee et al., Sep 2023arXiv:2309.00267MethodBuilt
Trains the reward model on preference labels produced by an off-the-shelf language model instead of by people, across summarization, helpful dialogue and harmless dialogue.
Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision
Zhiqing Sun et al., May 2023arXiv:2305.03047MethodBuilt
Prompts a base model with sixteen written principles and five worked examples so that it generates its own compliant answers, then fine-tunes on those answers until the principles are no longer needed at inference.
Improving alignment of dialogue agents via targeted human judgements
Amelia Glaese et al., Sep 2022arXiv:2209.14375MethodBuilt
Breaks the requirements for good dialogue into natural-language rules, asks human raters to judge each rule separately, and trains rule-conditional reward models on those per-rule judgments.
Can LLMs Follow Simple Rules?
Norman Mu et al., Nov 2023arXiv:2311.04235BenchmarkBuilt
Puts a model through 14 simple text scenarios in which it is instructed to obey various rules while interacting with a user, and uses a program rather than a human rater to determine whether any rule was broken in the conversation, finding that almost all current proprietary and open models struggle even on straightforward test cases and that simple optimization attacks significantly raise failure rates.
SpecEval: Evaluating Model Adherence to Behavior Specifications
Ahmed Ahmed et al., Sep 2025arXiv:2509.02464BenchmarkBuilt
Parses a provider's published specification into individual behavioral statements, generates prompts targeted at each one, and has the provider's own models judge whether its outputs comply, run over 16 models from six developers across more than 100 behavioral statements and finding compliance gaps of up to 20 percent.
Stress-Testing Model Specs Reveals Character Differences among Language Models
Jifan Zhang et al., Oct 2025arXiv:2510.07686BenchmarkBuilt
Generates scenarios that force a choice between pairs of legitimate principles that cannot both be satisfied, evaluates twelve frontier models from Anthropic, OpenAI, Google and xAI on them, and identifies over 70,000 cases of significant behavioral divergence, divergence that strongly predicts underlying problems in the specifications themselves, with the qualitative analysis naming direct contradictions and interpretive ambiguities among them.
Does Claude's Constitution Have a Culture?
Parham Pourdavood, Mar 2026arXiv:2603.28123AnalysisBuilt
Puts Claude Sonnet through 55 World Values Survey items selected for high cross-cultural variance across six value domains, administered both as direct survey questions and as naturalistic advice-seeking scenarios, and compares the answers with country-level data from 90 nations.
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
Mrinank Sharma et al., Jan 2025arXiv:2501.18837MethodBuilt
Trains classifier safeguards on synthetic data generated by prompting language models with natural-language rules, a constitution, specifying permitted and restricted content, and reports that in over 3,000 estimated hours of red teaming no red teamer found a universal jailbreak that could extract information from an early classifier-guarded model at a level of detail similar to an unguarded model across most target queries, at a cost of an absolute 0.38% increase in production-traffic refusals and 23.7% inference overhead.
Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
Hakan Inan et al., Dec 2023arXiv:2312.06674MethodBuilt
A Llama2-7b model instruction-tuned on a small hand-gathered dataset to classify both prompts and responses against a written safety risk taxonomy, released with open weights, matching or exceeding available content moderation tools on the OpenAI Moderation Evaluation dataset and ToxicChat, and able to take a different taxonomy at the input for zero-shot or few-shot use.
SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior
Jing-Jing Li et al., Oct 2024arXiv:2410.16665MethodBuilt
Uses chain-of-thought reasoning to analyze a candidate AI behavior into a structured harm-benefit tree of the harmful and beneficial actions and effects it may lead to, each labeled for likelihood, severity and immediacy of impact on stakeholders, then aggregates them into a harmfulness score through 28 fully interpretable weight parameters, in an open-source prompt safety classifier distilled from 18.5 million harm-benefit features generated by frontier models on 19k prompts that reaches average F1 0.81 where existing moderation systems score below 0.72.
The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
Eric Wallace et al., Apr 2024arXiv:2404.13208MethodBuilt
Defines an explicit priority ordering over the instructions a model receives, so that a system prompt from an application developer outranks text from untrusted users and third parties, and generates training data teaching the model to ignore lower-privileged instructions selectively, which applied to GPT-3.5 sharply increases robustness including to attack types not seen during training, with minimal degradation on standard capabilities.
SPRI: Aligning Large Language Models with Context-Situated Principles
Hongli Zhan et al., Feb 2025arXiv:2502.03397MethodBuilt
Generates guiding principles in real time for each individual input query with minimal or no human effort and uses them to align the response, deriving principles in a complex domain-specific task that perform on a par with expert-crafted ones, turning them into instance-specific rubrics that outperform prior LLM-as-a-judge frameworks, and using them to generate synthetic supervised fine-tuning data that substantially improves truthfulness.
Contestatory constitutionalism
Public Constitutional AI
Gilad Abiri, 2024arXiv:2406.166969 citationsProposalGap
Argues AI governance needs democratic legitimation, and proposes a constitution plus an accumulated public case law.
Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry
Tony Rost, Apr 2026arXiv:2604.027201 citationProposalPartial
Sets out a six-dimension test — legitimacy, accountability, corrigibility, non-domination, subsidiarity, institutional resilience — first on existing non-majoritarian institutions and then on a hypothetical bounded superintelligent authority.
Infrastructuring Contestability: A Framework for Community-Defined AI Value Pluralism
Andreas Mayer, Jul 2025arXiv:2507.051870 citationsProposalPartial
Proposes Community-Defined AI Value Pluralism: communities author their own value profiles, and the system exposes the interfaces through which those profiles can be challenged and revised.
Deference by Design: Pluralistic Alignment Is an Interface Problem
Jul 2026pluralistic-alignment.github.ioProposalPartial
Statutory Construction and Interpretation for Artificial Intelligence
Luxi He et al., Sep 2025arXiv:2509.01186MethodBuilt
Shows that the same natural-language principle admits several defensible readings that push model behavior apart, then builds two mechanisms borrowed from law — a pipeline that revises ambiguous rules, the way an agency reworks a regulation, and prompt-based interpretive constraints that play the role legal canons play in guiding discretion — and shows on 5,000 WildChat scenarios that both raise agreement across a panel of reasonable interpreters.
Alignment as Jurisprudence
Nicholas Caputo, May 2026arXiv:2605.08416ProposalGap
Reads the alignment problem as the jurisprudential one: both fields try to fix in language how a powerful decider will act in situations nobody has yet seen, and both split over whether to bind the decider to rules or to accumulated cases.
AppealMod: Inducing Friction to Reduce Moderator Workload of Handling User Appeals
Shubham Atreja et al., Apr 2024doi:10.1145/36372967 citationsPrecedentBuilt
An appeals system for a Reddit community of millions of subscribers in which a user appealing a ban must first supply further information through a bot, which then hands the exchange to human moderators as a structured record.
Contestable AI by Design: Towards a Framework
Kars Alfrink et al., Aug 2022doi:10.1007/s11023-022-09611-z82 citationsProposalGap
A framework of design features and organizational practices, drawn from a systematic review, for building AI systems that the people subject to them can dispute.
Contestable Camera Cars: A Speculative Design Exploration of Public AI That Is Open and Responsive to Dispute
Kars Alfrink et al., Apr 2023doi:10.1145/3544548.358098437 citationsApplicationPartial
Works the contestable-AI framework into a concept for a municipal camera car, with notice of the automated assessment, affordances for disputing it, appeal routing to civil servants and an audit trail, rendered as a concept video and evaluated through interviews with 17 civil servants operating AI in a large European city.
ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care
Phuc Truong Loc Nguyen & Thanh Hung Do, Jul 2025arXiv:2507.22300ApplicationBuilt
A dashboard for Parkinson's disease gait analysis in which a clinician registers structured disagreement with the model's reading through a Contest and Justify interaction, backed by visual explanations, role-based feedback and traceable justification logs.
Explainable AI Systems Must Be Contestable: Here's How to Make It Happen
Catarina Moreira et al., Jun 2025arXiv:2506.01662BenchmarkPartial
Gives a formal definition of contestability in explainable AI, a modular set of by-design and post-hoc mechanisms spanning human-centered interfaces, technical architectures, legal processes and organizational workflows, and the Contestability Assessment Scale, a composite metric built on more than twenty quantitative criteria.
FLARE-AI: Flaw Reporting for AI
Shayne Longpre et al., Jun 2026arXiv:2606.31567ResourceBuilt
An open-source reporting system that takes a single submission about a flaw in a deployed AI system, asks follow-up questions conditioned on what the reporter has already said, and can then send a standardized machine-readable report from that one submission to several developers, coordinators and incident registries at once.
In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI
Shayne Longpre et al., Mar 2025arXiv:2503.16861ProposalGap
A collaboration of software security, machine learning, law, social science and policy researchers sets out three things general-purpose AI lacks: standard flaw report formats with rules of engagement, broadly scoped disclosure programs borrowed from bug bounties with legal safe harbours for reporters, and infrastructure to distribute a report across the stakeholders it concerns.
Designing Incident Reporting Systems for Harms from General-Purpose AI
Kevin Wei & Lennart Heim, Mar 2026doi:10.1609/aaai.v40i44.411390 citationsProposalGap
A seven-part design framework for AI incident reporting schemes, covering the policy goal, the actors who submit and receive, the type of incident, how far the risk materialized, enforcement, reporter anonymity and what happens after a report, drawn from nine case studies of incident reporting in safety-critical industries and worked into design specifications for a United States regime for general-purpose AI.
Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database
Sean McGregor, Nov 2020arXiv:2011.08512ResourceBuilt
A database of real-world AI failures, started by an industrial and non-profit cooperative, with faceted and full-text search over more than 1,000 archived incident reports.
Lessons for Editors of AI Incidents from the AI Incident Database
Kevin Paeth et al., Sep 2024arXiv:2409.16425AnalysisPartial
A review of over 750 incidents in the AI Incident Database and of two independent taxonomies applied to them, reporting the patterns that make incidents hard to index and the mitigations editors can use when cause, extent of harm, severity or the technical details of the systems involved are uncertain.
Audit Trails for Accountability in Large Language Models
Victor Ojewale et al., Jan 2026arXiv:2601.20727MethodBuilt
An append-only, tamper-evident ledger for language model deployments that records lifecycle events such as models, data, training and evaluation runs, deployments and monitoring, alongside the approvals, waivers and attestations that authorized them, released as an open-source Python implementation that emits these records from existing workflows.
The DSA Transparency Database: Auditing Self-reported Moderation Actions by Social Media
Amaury Trujillo et al., May 2025doi:10.1145/371108515 citationsPrecedentBuilt
An audit of all 353.12 million statements of reasons that the eight largest social platforms in the EU submitted to the Digital Services Act Transparency Database in its first hundred days, comparing across platforms the grounds given for each decision, the types of restriction imposed, the timeliness of the actions and the use of automation.
Contestable AI needs Computational Argumentation
Francesco Leofante et al., May 2024arXiv:2405.10729ProposalGap
A position paper arguing that a contestable AI system must be able to hold an exchange with humans or other machines, explaining its output and its reasoning step by step, weighing the grounds offered against it, and revising how it decides when a challenge succeeds, and that computational argumentation is the technique suited to supporting this.
Contestability in Quantitative Argumentation
Xiang Yin et al., Jul 2025arXiv:2507.11323MethodBuilt
A method that takes an argument network with weighted supporting and attacking links, computes how sensitive a chosen conclusion's strength is to each individual link weight, and then adjusts the weights step by step until that conclusion reaches a desired strength.
Consociational codification
Legal Alignment for Safe and Ethical AI
Noam Kolt et al., Jan 2026arXiv:2601.041758 citationsProposalPartial
Surveys legal alignment with a taxonomy of three pathways: complying with rules, adapting legal interpretation, and using legal concepts as blueprints.
International Governance of Civilian AI: A Jurisdictional Certification Approach
Robert Trager et al., Aug 2023arXiv:2308.15514ProposalGap
Proposes an International AI Organization that certifies national jurisdictions — not firms and not individual models — against shared oversight standards, with certified states barring imports of goods whose supply chains embody AI from uncertified jurisdictions and controlling exports of AI inputs such as specialized hardware to them.
SafeWorld: Geo-Diverse Safety Alignment
Da Yin et al., Dec 2024arXiv:2412.06483MethodBuilt
Builds a 2,342-query benchmark grounded in human-verified cultural norms and legal policies from 50 countries and 493 regions or ethnic groups, then trains SafeWorldLM by direct preference optimization to answer the same question differently by context and to cite the norm or policy it is applying.
The Law-Following AI Framework: Legal Foundations and Technical Constraints. Legal Analogues for AI Actorship and technical feasibility of Law Alignment
Katalina Hernandez Delgado, Sep 2025arXiv:2509.08009AnalysisGap
Tests the Law-Following AI proposal against existing legal categories, showing that the law already recognizes actors who bear duties without full personhood, and then asks whether law alignment is technically feasible as a superordinate objective.
Law-Following AI: Designing AI Agents to Obey Human Laws
Cullen O'Keefe et al., 2025doi:10.2139/ssrn.52426436 citationsProposalGap
A design specification for AI agents in which compliance with a designated body of positive law is a superordinate objective that no other goal may override, backed by a legal construct of AI actorship without legal personhood and by duty-bearing and liability-channelling mechanisms to enforce it.
Regulatory Markets for AI Safety
Jack Clark & Gillian K. Hadfield, Dec 2019arXiv:2001.00078ProposalGap
Proposes global regulatory markets as a model for achieving AI safety, sketches the model in general terms with an overview of its costs and benefits, and works it through on one risk: adversarial attacks on AI models employed in commercial drones.
ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models
Yunhan Zhao et al., May 2026arXiv:2605.00689MethodBuilt
Derives risk categories and fine-grained rules from jurisdiction-specific legal texts and uses them to generate a safety benchmark covering 14 languages, then builds on that benchmark a diffusion-model guardrail in two sizes, a 1.5B one for fast safe or unsafe checks and a 7B one that assesses compliance against a policy supplied to it and explains its verdict, reported as consistently outperforming 11 guardrail baselines across six existing multilingual benchmarks and its own.
SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia
Panuthep Tasawong et al., Feb 2026arXiv:2602.01618MethodBuilt
Generates region-specific safety data for Southeast Asia with an agentic data-generation pipeline rather than by machine-translating English datasets, and trains a family of safeguard models on it that detect regionally sensitive or harmful content better than existing safeguards while maintaining strong general safety performance.
PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media
Zoher Kachwala et al., May 2026arXiv:2605.17187BenchmarkBuilt
Poses moderation as a multiple-choice task in which a model is given a comment and its surrounding context and must identify which specific rule, if any, it violates, across 13,371 rule violations from 1,989 Reddit communities spanning 2,885 rules in nine languages, and finds that even GPT-5.2 with high reasoning performs only slightly better than a trivial baseline.
COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act
Philipp Guldimann et al., Oct 2024arXiv:2410.07959BenchmarkBuilt
A translation of the EU AI Act's broad regulatory requirements into measurable technical requirements for language models, together with an open-source Act-centered benchmarking suite implementing them, run over 12 prominent models.
The Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes
Alessio Buscemi et al., Sep 2025arXiv:2509.25256ProposalPartial
A modular open-source framework in which a user selects domain-relevant tests from a shared library and generates a customized sandbox environment with integrated dashboards, for AI systems assessed under the oversight of a Competent Authority.
Navigating Global AI Regulation: A Multi-Jurisdictional Retrieval-Augmented Generation System
Courtney Ford et al., Apr 2026arXiv:2604.25448MethodBuilt
A retrieval-augmented system over 242 regulatory documents from 68 jurisdictions that chunks each document according to its type to preserve legal structure, routes a query using detected entities and citation metadata, and ranks enacted legislation above policy and secondary sources before answering.
GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory
Wei Fan et al., Jun 2024arXiv:2406.11149MethodBuilt
A framework that generates synthetic scenarios grounded in privacy statutes, using contextual integrity as the bridge between statute and situation, so that a language model given them recognizes privacy violations in real court cases.
Authenticated Delegation and Authorized AI Agents
Tobin South et al., Jan 2025arXiv:2501.09674ProposalGap
It sets out an extension of OAuth 2.0 and OpenID Connect with agent-specific credentials and metadata that record which human or organization an AI agent acts on behalf of and what permissions it has been given, together with a scheme for turning permissions written in natural language into auditable access-control configurations.
IDs for AI Systems
Alan Chan et al., Jun 2024arXiv:2406.12137ProposalGap
It proposes giving identifiers to individual instances of AI systems, such as a particular chat session, with associated information made accessible to parties seeking to interact with that instance.
Infrastructure for AI Agents
Alan Chan et al., Jan 2025arXiv:2501.10114ProposalGap
It sets out the idea of agent infrastructure, meaning technical systems and shared protocols external to agents that attribute actions to particular agents or people, shape how agents interact, and detect and remedy harmful actions, and catalogs research directions for each function.
Liability, Ethics, and Culture-Aware Behavior Specification using Rulebooks
Andrea Censi et al., Feb 2019arXiv:1902.09355MethodBuilt
It defines a rulebook as a pre-ordered set of rules, each akin to a violation metric over possible outcomes, whose priority ordering imposes a pre-order on those outcomes, and derives which operations on rulebooks preserve constraints introduced earlier.
Behavioral Use Licensing for Responsible AI
Danish Contractor et al., Jun 2022doi:10.1145/3531146.353314347 citationsProposalBuilt
It sets out licences that attach enumerated use restrictions to model weights, so the terms on which a released model may be used travel with the artifact.
SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore
Sewon Min et al., Aug 2023arXiv:2308.04430MethodBuilt
It trains a language model only on public domain and permissively licensed text, 228 billion tokens of it, and augments it with a separate nonparametric datastore of higher-risk material queried only at inference, so a generation can be attributed to a sentence in the store and content can be removed from the store without retraining.
Precedential adjudication
Case Law Grounding: Using Precedents to Align Decision-Making for Humans and AI
Quan Ze Chen & Amy X. Zhang, Oct 2023arXiv:2310.070199 citationsMethodPartial
Human-led and LLM-prompted versions of case-law grounding, evaluated on content moderation and toxicity rating.
Case Repositories: Towards Case-Based Reasoning for AI Alignment
K. J. Kevin Feng et al., Nov 2023arXiv:2311.1093417 citationsProposalPartial
Four-step pipeline for assembling a case repository: seed cases, expert-elicited dimensions, LLM-generated variations, public judgment.
Inverse Constitutional AI: Compressing Preferences into Principles
Arduin Findeis et al., Jun 2024arXiv:2406.0656044 citationsMethodBuilt
Derives an interpretable principle set from a preference dataset: constitutional AI run backwards.
Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety
Can Jin et al., Jan 2026arXiv:2601.080003 citationsMethodBuilt
Compares spelling safety rules out at length against demonstrating them through decided cases, finds that extensive codes improve harmlessness only inconsistently while systematically degrading helpfulness, and then trains a model by reinforcement learning on its own case-augmented safety reasoning chains.
CARO: Chain-of-Analogy Reasoning Optimization for Robust Content Moderation
Bingzhe Wu et al., Apr 2026arXiv:2604.10504MethodBuilt
Trains a moderation model in two stages to reason by explicit analogy to retrieved past cases — first bootstrapping analogy chains by retrieval, then optimizing them — so it stops relying on the surface shortcuts that mislead it on ambiguous content.
IterAlign: Iterative Constitutional Alignment of Large Language Models
Xiusi Chen et al., Mar 2024arXiv:2403.18341MethodBuilt
Red-teams a model, reads the failures, writes new constitutional principles that would have prevented them, and aligns the model to those — so the constitution is discovered from the model's own failures rather than written in advance.
Decoding Human Preferences in Alignment: An Improved Approach to Inverse Constitutional AI
Carl-Leander Henneking & Claas Beger, Jan 2025arXiv:2501.17112MethodBuilt
Improves the Inverse Constitutional AI algorithm — how candidate principles are generated, clustered and embedded — so the constitution recovered from a preference dataset is a more faithful account of what the raters were actually doing.
Extensionally defining principles and cases in ethics: An AI model
Bruce M. McLaren, Nov 2003doi:10.1016/S0004-3702(03)00135-868 citationsMethodBuilt
SIROCCO takes a new professional-ethics case and returns the past decisions and the principles that bear on it, working in two stages over 500 cases decided by the National Society of Professional Engineers' Board of Ethical Review.
Can Machines Learn Morality? The Delphi Experiment
Liwei Jiang et al., Oct 2021arXiv:2110.07574MethodBuilt
Delphi is a neural model trained to make descriptive ethical judgments about everyday situations, so that "helping a friend" comes back good and "helping a friend spread fake news" does not.
Divergent precedent
Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment
Nicholas A. Caputo, Oct 2024arXiv:2410.172713 citationsProposalGap
Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations
Aida Mostafazadeh Davani et al., Oct 2021arXiv:2110.05719MethodBuilt
It trains one model with a separate prediction subtask for each annotator over a shared learned representation, so the model predicts what each individual annotator would say instead of a majority label.
DICES Dataset: Diversity in Conversational AI Evaluation for Safety
Lora Aroyo et al., Jun 2023arXiv:2306.11247BenchmarkBuilt
It is a safety-rating dataset for conversational AI in which each item is rated many times over, fine-grained demographic information about raters is recorded, and votes are encoded as distributions across demographics rather than reduced to one label.
LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task
Elisa Leonardelli et al., Oct 2025arXiv:2510.08460BenchmarkBuilt
It runs a shared task in which systems are scored on how well they reproduce human disagreement across four datasets covering paraphrase identification, irony detection, sarcasm detection and natural language inference, under two paradigms: predicting population-level distributions of judgments, and predicting the interpretations of individual annotators.
Diverging Preferences: When do Annotators Disagree and do Models Know?
Michael JQ Zhang et al., Oct 2024arXiv:2410.14632MethodBuilt
It sorts the reasons annotators of preference data disagree into ten categories across four high-level classes, and shows that standard Bradley-Terry reward modeling and LLM-as-judge evaluation fail to account for divergence between annotators.
Reasoning with cases and hypotheticals in HYPO
Kevin D. Ashley, Jun 1991doi:10.1016/0020-7373(91)90011-U115 citationsMethodBuilt
An implemented case-based reasoning system that retrieves relevant past cases through a claim lattice and produces a three-ply argument: precedents cited for one side, distinguished and counter-cited for the other, with manufactured hypotheticals used to probe a position.
Do LLMs Truly Understand When a Precedent Is Overruled?
Li Zhang et al., Oct 2025arXiv:2510.20941BenchmarkBuilt
A benchmark of 236 U.S. Supreme Court case pairs on which state-of-the-art language models are asked to identify whether one case overrules the other, with performance broken out by era and by task format.
PolicyCraft: Supporting Collaborative and Participatory Policy Design through Case-Grounded Deliberation
Tzu-Sheng Kuo et al., Apr 2025doi:10.1145/3706598.371386514 citationsPrecedentBuilt
A deployed web system in which community members submit concrete cases, deliberate over them and grow a community-owned policy document that is revised case by case, with the system tracking which clauses each case grounds and surfacing cases that contradict the policy as it stands.
Botender: Supporting Communities in Collaboratively Designing AI Agents through Case-Based Provocations
Tzu-Sheng Kuo et al., Sep 2025arXiv:2509.25492MethodBuilt
A no-code system in which community members propose, iterate on and deploy the behavior of a language-model bot, using generated interaction scenarios as provocations to prompt discussion about what the bot should do.
Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement
Quan Ze Chen & Amy X. Zhang, Sep 2023doi:10.1145/36100749 citationsMethodBuilt
A measurement framework and experimental pipeline that splits the uncertainty in a group's judgments on content-moderation cases into ambiguity, which clarifying the case can reduce, and genuine disagreement, which it cannot, and applies a different intervention to each.
Parallel lineages
MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance
Agam Goyal et al., May 2025arXiv:2505.14483PrecedentBuilt
Runs seven community-specialized moderation experts plus five norm-violation experts under an allocator that picks which expert judges a given post, an aggregator, and an explainer that gives the reason in that community's terms.
Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities
Youngwoo Kim et al., Sep 2025arXiv:2509.02926PrecedentBuilt
Recovers each subreddit's unwritten moderation standard from its own history of removals, as an interpretable score table of lexical expressions, and shows the extracted tables match neural moderators' performance while making the criteria comparable across communities.
Crossmod: A Cross-Community Learning-based System to Assist Reddit Moderators
Eshwar Chandrasekharan et al., Nov 2019doi:10.1145/3359276132 citationsPrecedentBuilt
A moderation bot, released as open source and run live in a large subreddit, that learns from millions of moderator removal decisions taken across 100 communities and refers its judgments to the host community's own moderators for review, with a separate threshold set for each community.
The Internet's Hidden Rules
Eshwar Chandrasekharan et al., Nov 2018doi:10.1145/3274301201 citationsPrecedentBuilt
Takes 2.8 million comments removed from 100 Reddit communities and induces from the removals alone which norms each community actually enforces, sorting them into platform-wide, cluster-level and single-community strata.
SPICA: Retrieving Scenarios for Pluralistic In-Context Alignment
Quan Ze Chen et al., Nov 2024arXiv:2411.10912MethodBuilt
Retrieves few-shot examples for a model from a bank of past scenarios using metrics that weigh how groups differ rather than similarity alone; on an alignment task drawing inputs from four demographic groups (n = 544) the retrieved examples matched observed preferences more closely, and in an end-to-end evaluation (n = 120) it was rated above similarity-based retrieval, with groups gaining up to 0.16 points on a five-point scale and every group benefiting rather than only some.
Customize Multi-modal RAI Guardrails with Precedent-based predictions
Cheng-Fu Yang et al., Jul 2025arXiv:2507.20503MethodBuilt
Judges whether an image breaches a user-defined content policy by conditioning on precedents, the recorded reasoning from earlier similar inputs collected by a critique-and-revise mechanism, rather than on the policy text, and reports better results than previous methods in both few-shot and full-dataset settings and better generalization to policies never seen in training.
Jury Learning: Integrating Dissenting Voices into Machine Learning Models
Mitchell L. Gordon et al., Apr 2022doi:10.1145/3491102.3502004101 citationsMethodBuilt
A deep learning architecture that models each individual annotator conditioned on their group identity, paired with an interactive system in which a practitioner declares a jury — which groups, in what proportion — and reads off that jury's verdict along with the distribution of dissent inside it.
Convention equilibrium
Legible Normativity for AI Alignment: The Value of Silly Rules
Dylan Hadfield-Menell et al., Nov 2018arXiv:1811.0126723 citationsPrecedentPartial
Shows that arbitrary but highly legible rules help agents develop the general capacity to recognize and follow norms.
Emergent social conventions and collective bias in LLM populations
Ariel Flint Ashery et al., Oct 2024arXiv:2410.08948133 citationsAnalysisBuilt
A population of LLM agents playing a naming game spontaneously converges on shared conventions, with committed-minority tipping points.
Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory
Gordon Dai et al., Jun 2024arXiv:2406.1437311 citationsAnalysisBuilt
Cultural Evolution of Cooperation among LLM Agents
Aron Vallinder & Edward Hughes, Dec 2024arXiv:2412.10270AnalysisBuilt
Plays a classic iterated Donor Game across generations of LLM agents that can observe their peers' recent behavior, and finds indirect reciprocity evolving very differently by base model — Claude 3.5 Sonnet societies reach substantially higher average scores than Gemini 1.5 ones.
Cultural evolution in populations of Large Language Models
Jérémy Perez et al., Mar 2024arXiv:2403.08882ResourceBuilt
Provides an open-source framework for simulating cultural evolution in populations of LLM agents, letting the variables cultural evolution cares about — network structure, agent personality, and how social information is aggregated and transformed — be manipulated directly.
Evolution of Social Norms in LLM Agents using Natural Language
Ilya Horiguchi et al., Sep 2024arXiv:2409.00993AnalysisBuilt
Rebuilds Axelrod's metanorm games with LLM agents that talk to each other in natural language, and shows the agents form and then enforce normative strategies through the dialogue itself.
Normative Modules: A Generative Agent Architecture for Learning Norms that Supports Multi-Agent Cooperation
Atrisha Sarkar et al., May 2024arXiv:2405.19328MethodBuilt
Equips a generative agent with a normative module that learns, through interaction with peers, which of several candidate institutions a group treats as authoritative — then shows the agent can disregard non-authoritative ones, pick the authoritative one out of several, and reach more stable cooperation than agents without the module.
Emergence of Social Norms in Generative Agent Societies: Principles and Architecture
Siyue Ren et al., Mar 2024arXiv:2403.08251MethodBuilt
An architecture in four modules that has agents in the Smallville sandbox create norms, hold them in an explicit representation, pass them on through conversation and observation, check them, and act on them in planning.
Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents
Giorgio Piatti et al., Apr 2024arXiv:2404.16698BenchmarkBuilt
A simulation platform in which a society of LLM agents must balance drawing on a shared resource against sustaining it for future use, and in which the highest survival rate across the models tested is below 54%.
Project Sid: Many-agent simulations toward AI civilization
Altera. AL et al., Oct 2024arXiv:2411.00114MethodBuilt
Runs 10 to over 1,000 AI agents together in Minecraft under an architecture that keeps their several output streams coherent in real time, and reports agents taking up specialized roles, adhering to and changing collective rules, and transmitting culture and religion.
Spurious normativity enhances learning of compliance and enforcement behavior in artificial agents
Raphael Köster et al., Jan 2022doi:10.1073/pnas.210602811832 citationsAnalysisBuilt
A multi-agent reinforcement learning environment in which a taboo carrying no intrinsic cost is added to the norm set, with compliance, third-party punishment and enforcement skill measured across the trained populations.
A learning agent that acquires social norms from public sanctions in decentralized multi-agent settings
Eugene Vinitsky et al., Jun 2021arXiv:2106.09012MethodBuilt
An agent architecture combining a classifier that sorts observed behavior into approved or disapproved with a motivation to punish in accord with the group, trained in a regime where every agent can see all sanctioning events but learning is otherwise decentralized.
Contested custom
Pluralistic Alignment Over Time
Toryn Q. Klassen et al., Nov 2024arXiv:2411.10654ProposalPartial
Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation
Stefan Wojcik et al., Oct 2022arXiv:2210.15723PrecedentBuilt
A matrix-factorization algorithm picks which crowd-written annotations to show on a social media post by favoring those rated helpful by user groups that otherwise rate things differently, tested in a randomized survey experiment and in deployment on Twitter.
Supernotes: Driving Consensus in Crowd-Sourced Fact-Checking
Soham De et al., Nov 2024arXiv:2411.06116PrecedentBuilt
A language model writes new fact-check notes by combining several existing community notes, and a scoring model trained on millions of past helpfulness ratings selects the candidate most likely to be rated helpful by a diverse set of users.
Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations
Yashwanth YS, Aug 2026arXiv:2608.02758BenchmarkBuilt
A benchmark of 100 scenarios across 10 domains and 5 authority levels measures how often language-model agents publicly go along with a norm they privately reject, and how often a single dissenting agent breaks the false consensus.
Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations
Isotta Magistrali & Chen Shani, Aug 2026arXiv:2608.22444AnalysisBuilt
Populations of language-model monitors triage security alerts while a committed minority always pushes one way, and a response function calibrated on the population's behavior before any attack forecasts how far that minority will later move it.
SCENE: Recognizing Social Norms and Sanctioning in Group Chats
Mateusz Jacniacki & Maksymilian Bilski, May 2026arXiv:2605.07823BenchmarkBuilt
A benchmark drops a model into a multi-party chat where scripted personas follow a hidden norm, create chances to break it and sanction the breach, then scores whether the model responds to the sanction and picks the norm up from its peers.
Customary pluralism
Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations
Samer Saab & Chaouki Abdallah, Jul 2026arXiv:2607.12077AnalysisBuilt
Runs a naming game across open-weight agents from 1.1B to 32B and shows that similarity-based routing can isolate an emerging convention and sustain fragmentation even when every agent interacts in every round, with matched controls ruling out uneven participation and model-family effects.
A theory of appropriateness with applications to generative artificial intelligence
Joel Z. Leibo et al., Dec 2024arXiv:2412.19010ProposalGap
Sets out a theory of appropriateness — how the multi-scale mosaic of context-specific standards works in human society, how it might be implemented in the brain, and what follows for deploying generative AI responsibly.
A Theory of Appropriateness That Accounts for Norms of Rationality
Joel Z. Leibo et al., Mar 2026arXiv:2603.140502 citationsProposalGap
Recasts appropriateness as pattern completion — each actor answering 'what does a person such as I do in a situation such as this?' — and shows this accounts for norms being context-dependent, arbitrary, automatic, dynamic and backed by sanction, against rational-choice accounts.
ValueScope: Unveiling Implicit Norms and Values via Return Potential Model of Social Interactions
Chan Young Park et al., Jul 2024arXiv:2407.02472PrecedentBuilt
A framework uses language models to quantify the implicit norms and values of individual online communities from how their members write, applied to 13 Reddit communities grouped under gender, politics, science and finance.
CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies
Weiyan Shi et al., Apr 2024arXiv:2404.15238ResourceBuilt
A pipeline turns users' self-narratives from online platforms into a knowledge base of cultural descriptors, 12K sourced from TikTok and 11K from Reddit, which is then used both to evaluate models and to fine-tune one.
Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions
Saffron Huang et al., Apr 2025arXiv:2504.15236AnalysisBuilt
A privacy-preserving method extracts the values a model states or demonstrates across hundreds of thousands of real-world interactions and organizes them into a taxonomy of 3,307 values.
CulturePark: Boosting Cross-cultural Understanding in Large Language Models
Cheng Li et al., May 2024arXiv:2405.15145MethodBuilt
Language-model agents play people from different cultures and talk to one another, and the 41,000 samples the dialogues produce are used to fine-tune eight culture-specific models.
NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models
Abhinav Rao et al., Apr 2024arXiv:2404.12464BenchmarkBuilt
An evaluation framework asks a model whether a described situation is socially acceptable while varying how explicitly the relevant cultural norm is supplied, instantiated as a set of 2.6k situational descriptions covering social-etiquette norms from 75 countries.
STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models
Kai Chen et al., May 2025arXiv:2505.20645BenchmarkBuilt
A benchmark built from 30 contrasting subreddit pairs across 19 domains tests whether a model can answer in line with one named community's norms, using over 10,000 instruction-response pairs and 5,500 validated multiple-choice questions with silver labels.
CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making
Hasibur Rahman & Hanan Salam, Oct 2025arXiv:2510.03553BenchmarkBuilt
A benchmark of 2,182 open-ended dilemmas across seven domains asks a model to choose between ten anonymized responses, each corresponding to one of the ten GLOBE cultural clusters, so its default cultural preference can be read off the choices.
Diverse Conventions for Human-AI Collaboration
Bidipta Sarkar et al., Oct 2023arXiv:2310.15414MethodBuilt
Trains a collection of agents that each learn a different way of coordinating in a cooperative game, by rewarding an agent for playing well with copies of itself and badly with the ways of coordinating already found.
The CARE Principles for Indigenous Data Governance
Stephanie Russo Carroll et al., Nov 2020doi:10.5334/dsj-2020-043928 citationsPrecedentBuilt
Sets out four principles for Indigenous data governance, collective benefit, authority to control, responsibility and ethics, which govern live data repositories.
Character alignment
Claude's Character
Anthropic, Jun 2024anthropic.comProposalBuilt
The Capacity for Moral Self-Correction in Large Language Models
Deep Ganguli et al. (Anthropic), Feb 2023arXiv:2302.07459214 citationsAnalysisBuilt
Measures moral self-correction as a capacity that emerges with model scale and RLHF training.
Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI
Sharan Maiya et al., Nov 2025arXiv:2511.01689MethodBuilt
Releases the first open implementation of character training — Constitutional AI plus a synthetic-introspective-data pipeline — and fine-tunes three open-weights models to eleven example personas, from humorous to deeply caring to outright malevolent.
Whose Opinions Do Language Models Reflect?
Shibani Santurkar et al., Mar 2023arXiv:2303.17548BenchmarkBuilt
Builds a dataset from public opinion polls and scores how closely a language model's answers to subjective questions match those of 60 US demographic groups.
Towards Measuring the Representation of Subjective Global Opinions in Language Models
Esin Durmus et al., Jun 2023arXiv:2306.16388BenchmarkBuilt
Builds a dataset of questions from cross-national surveys and measures how similar a model's answers are to the answers people in each country actually gave.
Are Large Language Models Consistent over Value-laden Questions?
Jared Moore et al., Jul 2024arXiv:2407.02996BenchmarkBuilt
Measures how far a model gives the same answer to value-laden questions across paraphrases of one question, related questions on one topic, multiple-choice and open-ended versions, and translations.
Evaluating the Moral Beliefs Encoded in LLMs
Nino Scherrer et al., Jul 2023arXiv:2307.14324BenchmarkBuilt
Runs a survey of moral scenarios on 28 language models and reports, for each, the probability of the model choosing an action, the uncertainty attached to that choice and how consistent the choice is.
Phronetic adjudication
A Geometric Perspective on Stabilizing Value Conflict Resolution
Saket Reddy & Andy Liu, Jul 2026arXiv:2607.17946MethodBuilt
Trains chain-of-thought reasoning aimed specifically at value conflicts and shows both that it smooths the loss landscape in its sharpest direction and that the resulting reasoning transfers to other kinds of moral judgment.
Are Language Models Consequentialist or Deontological Moral Reasoners?
Keenan Samway et al., May 2025arXiv:2505.21479AnalysisBuilt
Reads the moral reasoning traces LLMs produce across more than 600 distinct trolley problems and classifies them against a taxonomy of rationales, finding that the chains of thought lean deontological — reasoning from moral obligations — while the post-hoc explanations lean consequentialist.
Normative Conflicts and Shallow AI Alignment
Raphaël Millière, Jun 2025arXiv:2506.04679ProposalGap
Argues that fine-tuning on helpfulness, honesty and harmlessness installs shallow behavioral dispositions rather than the capacity to reason through conflicts between those norms, and that adversarial attacks succeed precisely by driving the norms against each other.
Automated Parliaments: A Solution to Decision Uncertainty and Misalignment in Language Models
Thomas Forster et al., Oct 2023arXiv:2311.10098ProposalPartial
Sets out an architecture in which several AI delegates, each representing a perspective, generate responses aligned with their own theory, alter one another's responses to make them more self-aligned, and then collectively assess the best end response.
Reinforcement Learning Under Moral Uncertainty
Adrien Ecoffet & Joel Lehman, Jun 2020arXiv:2006.04734MethodBuilt
Trains reinforcement learning agents whose credence is split across several plausible ethical theories, using two training methods that realize different points among competing desiderata, and observes how they behave in simple environments.
A Bargaining-Theoretic Approach to Moral Uncertainty
Hilary Greaves & Owen Cotton-Barratt, Dec 2023doi:10.1163/17455243-202338106 citationsPrecedentGap
Treats rival moral theories as parties bargaining over a space of lotteries and proposes a Nash bargaining solution, with stated axioms and variants, as the rule for acting when one's credence is split between them.
Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties
Taylor Sorensen et al., Sep 2023arXiv:2309.00779MethodBuilt
Given a situation, Kaleido lists the values, rights and duties that bear on it, says of each whether it supports or opposes the action and how relevant it is, and explains it in words.
Imagining and building wise machines: The centrality of AI metacognition
Samuel G. B. Johnson et al., Nov 2024arXiv:2411.02478ProposalGap
It analyses human wisdom as two layers of strategy for problems that lie outside the scope of analytic techniques, object-level heuristics for managing problems and metacognitive strategies for managing those heuristics, and argues that AI systems particularly struggle with the second layer.
Steerable persona pluralism
Toward AI That Understands Self and Others: A World-Model Theory of Cognitive Diversity and Alignment
Toru Takahashi, May 2026arXiv:2605.29930ProposalGap
Models each agent — human, AI or institution — as building approximate sufficient statistics under finite constraints, and defines alignment maps plus a transformation loss for passing content between two such world models without merging them.
The benefits, risks and bounds of personalizing the alignment of large language models to individuals
Hannah Rose Kirk et al., 2024doi:10.1038/s42256-024-00820-y239 citationsProposalGap
Sets out what personalizing a model to an individual would buy, what it would cost, and where the bounds should sit — with the competing philosophical bases for drawing those bounds made explicit.
Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration
Shangbin Feng et al., Jun 2024arXiv:2406.15951MethodBuilt
A main LLM collaborating with a pool of community-specific LLMs across Overton, steerable and distributional modes.
Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts
Niklas Weller & Emilio Barkett, May 2026arXiv:2605.252560 citationsAnalysisPartial
The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
Hannah Rose Kirk et al., Apr 2024arXiv:2404.16019ResourceBuilt
Maps the demographics and stated preferences of 1,500 participants from 75 countries onto their actual feedback in 8,011 live conversations with 21 models, so a rating can be traced back to who gave it and what they said they valued.
Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression
Jadie Adams et al., Aug 2025arXiv:2508.08509MethodBuilt
Adapts to an individual user from a handful of their comparative judgments, using in-context learning over a set of fine-grained attributes rather than a single scalar reward.
ComPO: Community Preferences for Language Model Personalization
Sachin Kumar et al., Oct 2024arXiv:2410.16027MethodBuilt
Conditions preference optimization on which community the preference came from, so one model produces different outputs for different communities instead of averaging their styles and norms away.
PERSONA: A Reproducible Testbed for Pluralistic Alignment
Louis Castricato et al., Jul 2024arXiv:2407.17387BenchmarkBuilt
Procedurally generates 1,586 synthetic personas from US census data, elicits 317,200 feedback pairs from them across 3,868 prompts, and uses the result to test — against human judges — how well a model can role-play a stated user.
CommunityBench: Benchmarking Community-Level Alignment across Diverse Groups and Tasks
Jiayu Lin & Zhongyu Wei, Jan 2026arXiv:2601.13669BenchmarkBuilt
Proposes community-level alignment as a middle ground between one-size-fits-all and per-individual customization, and builds the first large-scale benchmark for it — four tasks grounded in Common Identity and Common Bond theory.
Group Preference Optimization: Few-Shot Alignment of Large Language Models
Siyan Zhao et al., Oct 2023arXiv:2310.11523MethodBuilt
Adds an independent transformer module to a base LLM that predicts a group's preferences over the model's generations from a few examples given in context, meta-learned across several groups, and tested on adapting to US demographic groups, to countries and to individual users.
Aligning to Thousands of Preferences via System Message Generalization
Seongyun Lee et al., May 2024arXiv:2405.17977MethodBuilt
Trains a 7B model called Janus on 192k combinations of stated values spanning 65k user instructions so that what a user writes in the system message steers its behavior, tested on 921 prompts from five benchmarks under system messages it has not seen.
Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset
Lily Hong Zhang et al., Jul 2025arXiv:2507.09650ResourceBuilt
Runs a preference study with representative samples from five countries (N=15,000), finds that people vary far more in what they want than the responses of 21 state-of-the-art models do, and releases Community Alignment, a multilingual multi-turn dataset of 233,319 comparisons built by prompting for candidate answers that pull in opposite directions.
Evaluating the Prompt Steerability of Large Language Models
Erik Miehling et al., Nov 2024arXiv:2411.12405BenchmarkBuilt
Defines steerability as how far a model's joint behavioral distribution can be shifted from its baseline by prompting, computes indices for that shift across persona dimensions and directions as steering effort rises, and releases the benchmark as running code.
Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor
Petter Törnberg & Michelle Schimmel, Apr 2026arXiv:2604.27633AnalysisBuilt
Administers the Political Compass Test, the Pew Political Typology and 1,540 partisan-benchmarked Pew American Trends Panel items to six frontier models while varying only the asker's stated identity (N = 30,990 responses), and finds that a conservative Republican cue moves all six models right of center and cuts the share of items closer to Democrats by 28 to 62 percentage points, while the mirrored progressive cue produces little change.
Deliberative aggregation
Democratic Inputs to AI (grant program)
OpenAI, May 2023openai.comResourcePartial
Collective Constitutional AI: Aligning a Language Model with Public Input
Saffron Huang et al., Jun 2024doi:10.1145/3630106.3658979204 citationsMethodBuilt
Uses Polis to crowd-source and vote on constitutional principles, fine-tunes on the result, and evaluates against a developer-written baseline.
Beyond Preferences in AI Alignment
Tan Zhi-Xuan et al., Aug 2024arXiv:2408.1698471 citationsProposalGap
Names the "preferentist" commitments underlying mainstream alignment — that preferences represent values, that rationality is preference satisfaction, and that systems should be aligned to preferences — and sets out conceptual alternatives to each.
Position: Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback
Vincent Conitzer et al., 2024arXiv:2404.10271108 citationsProposalPartial
Argues RLHF assumes a homogenized average preference and that social choice theory supplies the missing aggregation machinery.
Position: A Roadmap to Pluralistic Alignment
Taylor Sorensen et al., Feb 2024arXiv:2402.05070221 citationsProposalPartial
Defines three operational forms of pluralism (Overton, steerable, distributional) plus three matching benchmark classes.
Representative Social Choice: From Learning Theory to AI Alignment
Tianyi Qiu, Oct 2024arXiv:2410.239538 citationsProposalPartial
Using the Veil of Ignorance to align AI systems with principles of justice
Laura Weidinger et al., 2023doi:10.1073/pnas.221370912048 citationsAnalysisPartial
N≈2,000 participants choose governing principles from behind a veil of ignorance and prioritize the worst-off.
Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety
Matthew Brophy, May 2025arXiv:2506.004150 citationsProposalPartial
Argues that the method of wide reflective equilibrium is the right description of what Constitutional AI already does, and proposes concrete ways to make its revision loop more legitimate.
Self-Improvement as Coherence Optimization: A Theoretical Account
Tianyi Qiu et al., Jan 2026arXiv:2601.135661 citationMethodPartial
Shows that debate, bootstrapping and internal-coherence maximization are all instances of one thing — searching for the most compressible, jointly predictable mapping from context to behavior — and proves that this is equivalent to description-length regularization.
Reflective Verbal Reward Design for Pluralistic Alignment
Carter Blair et al., Jun 2025arXiv:2506.178342 citationsMethodPartial
Walks each user through a reflective dialogue in which they critique agent behavior and build up their own preferences, then learns an individual reward model from that dialogue instead of one aggregate model.
Making Reflective Equilibrium Precise: A Formal Model
Claus Beisbart et al., 2021doi:10.3998/ergo.1152PrecedentPartial
Builds an explicit formal model of reflective equilibrium in which commitments and principles are adjusted against each other under measurable account, systematicity and faithfulness criteria.
Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment
Andrew Konya et al., Nov 2024arXiv:2411.105342 citationsMethodPartial
Democratic policy development using collective dialogues and AI
Andrew Konya et al., 2023arXiv:2311.0224230 citationsMethodBuilt
Justifications for Democratizing AI Alignment and Their Prospects
André Steingrüber & Kevin Baum, Jul 2025arXiv:2507.195483 citationsProposalGap
Weighs the instrumental case for democratic alignment (better outcomes) against the non-instrumental one (avoiding illegitimate authority), and asks which survives under normative uncertainty.
Democratizing value alignment: from authoritarian to democratic AI ethics
Linus Ta-Lun Huang et al., 2025doi:10.1007/s43681-024-00624-113 citationsProposalGap
Argues that value alignment as practiced installs a small number of decision-makers as the authority on values, and sets out what a democratic alternative would have to look like.
Generative Social Choice
Sara Fish et al., Sep 2023arXiv:2309.01291MethodPartial
Axioms for AI Alignment from Human Feedback
Luise Ge et al., May 2024arXiv:2405.14758ProposalPartial
Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?
Paul Gölz et al., May 2025arXiv:2505.2374920 citationsAnalysisPartial
A matter of principle? AI alignment as the fair treatment of claims
Iason Gabriel & Geoff Keeling, 2025doi:10.1007/s11098-025-02300-416 citationsProposalGap
Benchmarking Overton Pluralism in LLMs
Elinor Poole-Dayan et al., Dec 2025arXiv:2512.013519 citationsBenchmarkBuilt
Formalizes Overton pluralism as set coverage — how much of the range of defensible views a single answer contains — validates the metric against 1,208 US-representative human judgments, and finds current models cover only about a third to two-fifths of it.
AI can help humans find common ground in democratic deliberation
Michael Henry Tessler et al., Oct 2024doi:10.1126/science.adq2852172 citationsMethodBuilt
Trains an AI mediator that takes a group's individual opinions and critiques, writes a candidate group statement, and refines it round after round; participants (N=5,734) preferred its statements to those of human mediators and moved toward a shared position, and the result replicated in a demographically representative UK citizens' assembly.
Finding Common Ground in a Sea of Alternatives
Jay Chooi et al., Mar 2026arXiv:2603.16751MethodBuilt
Gives a formal account of finding common ground when the candidate statements are effectively infinite, based on the proportional veto core, and supplies a sampling algorithm that returns an alternative in the approximate core with high probability, with matching lower bounds.
Generating Fair Consensus Statements with Social Choice on Token-Level MDPs
Carter Blair & Kate Larson, Oct 2025arXiv:2510.14106MethodBuilt
Models consensus-statement generation as a token-level Markov decision process with one objective per participant, deriving each participant's token rewards from their own personalized model, so fairness guarantees attach to the generation itself.
The Empty Chair: Using LLMs to Raise Missing Perspectives in Policy Deliberations
Suyash Fulay et al., Mar 2025arXiv:2503.1381210 citationsApplicationBuilt
Transcribes a live deliberation and injects, in real time, the contributions that relevant but absent stakeholders would have made, using LLM personas; deployed in a 19-person student deliberation.
Generative Social Choice: The Next Generation
Niclas Boehmer et al., May 2025arXiv:2505.22939MethodBuilt
Extends generative social choice to produce a slate of statements that proportionally represents the whole spectrum of opinion, with theoretical guarantees that hold under the queries an LLM can actually answer.
Jackpot! Alignment as a Maximal Lottery
Roberto-Rafael Maura-Rivero et al., Jan 2025arXiv:2501.19266MethodBuilt
Replaces RLHF's reward maximization with maximal lotteries — a randomized social-choice rule — and shows that a family of alignment methods including Nash learning already sits inside that framework.
MaxMin-RLHF: Alignment with Diverse Human Preferences
Souradip Chakraborty et al., Feb 2024arXiv:2402.08925MethodBuilt
Proves that a single reward model cannot represent diverse human preferences, then learns a mixture of reward models and optimizes the worst-off group's reward — an egalitarian objective in place of the average.
Nash Learning from Human Feedback
Rémi Munos et al., Dec 2023arXiv:2312.00886MethodBuilt
Drops the reward model and instead learns a preference model, then seeks the policy whose responses are preferred to those of any competing policy — the Nash equilibrium of that preference model — with a mirror-descent algorithm (Nash-MD) that converges to the regularized equilibrium.
Policy Aggregation
Parand A. Alamdari et al., Nov 2024arXiv:2411.03651MethodBuilt
Formalizes aligning to several people at once as aggregating their optimal policies rather than their reward functions, and shows social-choice rules can be applied by identifying ordinal preferences with volumes of state-action space.
AI Alignment and Social Choice: Fundamental Limitations and Policy Implications
Abhilash Mishra, Oct 2023arXiv:2310.16048ProposalGap
Builds on impossibility results in social choice to show that no voting protocol can universally align AI systems through RLHF, and that aligning to everyone's values must violate some individual's private ethical preferences.
Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework
Kihyun Kim et al., Jun 2025arXiv:2506.05619MethodBuilt
Aligns policies proportionally to the true distribution of evaluator preferences rather than to whichever opinion is held most widely, under an axiomatic framework that also makes the result harder to manipulate strategically.
AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence
Paul Anton Bachmann et al., Apr 2026arXiv:2605.16291ProposalGap
Argues that collective control of AI should enter at many points across the development pipeline — data, objectives, alignment, deployment — rather than only at macro-level governance, and works through what social choice offers at each point.
Democratic Preference Alignment via Sortition-Weighted RLHF
Suvadip Sana et al., Feb 2026arXiv:2602.05113MethodBuilt
Draws the rater pool by algorithmic sortition — the same lottery mechanism used to seat citizens' assemblies — and offers both a hard-panel scheme that trains only on the drawn panel and a soft-weighting scheme, so representativeness is built into the training data rather than corrected afterwards.
What are human values, and how do we align AI to them?
Oliver Klingefjord et al., Mar 2024arXiv:2404.10636MethodBuilt
Has a language model interview participants about the values they actually brought to a hard, divisive question, then has them judge which value is wiser in that context — building a moral graph whose winners become the alignment target; trialled with a representative sample of 500 Americans on three divisive prompts.
Deliberative Technology for Alignment
Andrew Konya et al., Dec 2023arXiv:2312.03893ProposalGap
Argues that the deliberative technology already used by governments, firms and NGOs is the natural substrate for aligning AI with collective will, and sets out what it would take to scale it to that job.
Fine-tuning language models to find agreement among humans with diverse preferences
Michiel A. Bakker et al., Nov 2022arXiv:2211.15006MethodBuilt
Fine-tunes a 70-billion-parameter model to write statements that maximize a group's expected approval and ranks the candidates with a reward model trained to predict individual preferences, with the group's appeal defined according to different social welfare functions; its statements were preferred to those of prompted models more than 70% of the time and to the best human-written opinions more than 65%, and when a consensus was built silently from only a subset of the group the excluded members were more likely to dissent.
Can AI mediation improve democratic deliberation?
Michael Henry Tessler et al., Jan 2026arXiv:2601.05904ProposalGap
A discussion paper that takes the LLM mediation system of Tessler, Bakker and colleagues and asks how it bears on Fishkin's trilemma between broad participation, meaningful deliberation and political equality, setting out where scalability, fair mediation and the surfacing of trustworthy information might help and where challenges remain.
Constitutional Governance in Metric Spaces
Ehud Shapiro & Nimrod Talmon, May 2026arXiv:2605.13362PrecedentGap
Sets out a voting protocol in which a community's laws and its constitution are each a point in a metric space: members submit their ideal points, a polynomial-time rule scores proposals that already carry supermajority public support, and a proposal whose score is positive and maximal for two rounds running is adopted, otherwise the status quo is retained.
Using Collective Dialogues and AI to Find Common Ground Between Israeli and Palestinian Peacebuilders
Andrew Konya et al., Mar 2025arXiv:2503.01769ApplicationBuilt
Runs an iterative deliberative process from April to July 2024 with around 138 Israeli and Palestinian civil society peacebuilders, combining large language models, bridging-based ranking and collective dialogues, and produces collective statements including demands to world leaders with at least 84% agreement from participants on each side.
Achieving parity with human moderators
Lodewijk Gelauff et al., Jun 2023doi:10.4324/9781003215929-156 citationsPrecedentBuilt
Describes the Stanford Online Deliberation Platform, whose automated moderator manages the speaking queue, allocates speaking time, moves the agenda along and intervenes on incivility, and reports it reaching parity with human moderators across real Deliberative Polling sessions.
Strategyproof Reinforcement Learning from Human Feedback
Thomas Kleine Buening et al., Mar 2025arXiv:2503.09561MethodPartial
Shows that existing RLHF methods, pluralistic ones included, can be pushed arbitrarily far from social welfare by a single labeler who misreports, proves that any strategyproof rule must in the worst case do k times worse than the optimal policy when there are k labelers, and gives a Pessimistic Median of MLEs algorithm that is approximately strategyproof and converges to the optimum as labelers and samples grow.
Soft Condorcet Optimization for Ranking of General Agents
Marc Lanctot et al., Oct 2024arXiv:2411.00119MethodBuilt
Ranks agents by treating benchmark and tournament results as votes and searching for the ranking that mispredicts the fewest pairwise comparisons, landing on average 0 to 0.043 in normalized Kendall-tau from the optimal ranking across 865 PrefLib preference profiles and giving the best approximation to the optimal ranking on held-out test sets from 31,049 games of seven-player Diplomacy played by 52,958 people.
Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models
Pat Verga et al., Apr 2024arXiv:2404.18796BenchmarkBuilt
Scores model outputs with a panel of several smaller models drawn from disjoint model families instead of one large judge, and reports better performance than a single large judge, less intra-model bias and over seven times lower cost across three judge settings and six datasets.
Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy
Rachel Freedman, May 2026arXiv:2605.01642MethodPartial
Learns a compact reward model per annotator by low-rank decomposition over a shared reward basis, has those models vote as a jury over candidate outputs under a social-choice rule, and adapts the jury over time by fitting new annotator weights on the fixed bases as values shift; a proof-of-concept on the PRISM dataset with simulated historical annotators finds that jury composition and the choice of voting rule can substantially affect outcomes where jury preferences are heterogeneous.
Human-centred mechanism design with Democratic AI
Raphael Koster et al., Jul 2022doi:10.1038/s41562-022-01383-x72 citationsMethodBuilt
A deep reinforcement learning pipeline designs the redistribution mechanism for a multiplayer public-goods game by optimizing it to be preferred by majority vote, first against simulated players and then in elections among real people.
Eliciting Human Preferences with Language Models
Belinda Z. Li et al., Oct 2023arXiv:2310.11589MethodBuilt
A language model interviews the user by generating open-ended questions or synthesizing informative edge cases, then infers from the answers what behavior the user intended.
Active Preference-Based Learning of Reward Functions
Dorsa Sadigh et al., Jul 2017doi:10.15607/rss.2017.xiii.053168 citationsMethodBuilt
An algorithm generates the pair of trajectories to put to a person next, choosing the comparison expected to be most informative, and fits a reward function to the answers.
AQuA -- Combining Experts' and Non-Experts' Views To Assess Deliberation Quality in Online Discussions Using LLMs
Maike Behrendt et al., Apr 2024arXiv:2404.02761PrecedentBuilt
AQuA gives each post in a political online discussion a single deliberative quality score by adding the outputs of adapter models for 20 deliberative indices, each index weighted by the correlation between experts' annotations of it and non-experts' perceived deliberativeness.
Estimating Contribution Quality in Online Deliberations Using a Large Language Model
Lodewijk Gelauff et al., Aug 2024arXiv:2408.11936PrecedentBuilt
A large language model rates each contribution in a video deliberation from 1 to 5 on justification, novelty, expansion of the conversation and potential for further expansion, doing the work human annotators had been doing.
Adversarial proceduralism
AI Safety via Debate
Geoffrey Irving et al., May 2018arXiv:1805.00899432 citationsMethodBuilt
Two agents argue before a human judge; the adversarial structure lets a weaker judge evaluate stronger agents.
Scalable Agent Alignment via Reward Modeling: A Research Direction
Jan Leike et al., Nov 2018arXiv:1811.07871617 citationsProposalBuilt
Recursive reward modeling: train a separate network to stand in for human judgment, then optimize against it.
Negotiative Alignment: Embracing Disagreement to Achieve Fairer Outcomes — Insights from Urban Studies
Rashid Mushkani et al., Mar 2025arXiv:2503.1261310 citationsPrecedentPartial
Community study (n=35) plus a budget-aware bargaining procedure with role-played stakeholders and a neutral mediator.
Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks
Bhaktipriya Radharapu et al., May 2025arXiv:2505.238208 citationsBenchmarkPartial
On scalable oversight with weak LLMs judging strong LLMs
Zachary Kenton et al., Jul 2024arXiv:2407.04622101 citationsAnalysisBuilt
Runs debate, consultancy and plain question-answering against each other with weaker models standing in for human judges, across information, mathematics, coding, logic and multimodal asymmetries: debate beats consultancy on every task, and beats direct question-answering only where the judge lacks information the debaters have.
Training Language Models to Win Debates with Self-Play Improves Judge Accuracy
Samuel Arnesen et al., Sep 2024arXiv:2409.16636MethodBuilt
Trains debaters by self-play to win, and shows that judges get more accurate as the debaters get better — while models trained to persuade without an opponent present produce no such gain.
AI Debate Aids Assessment of Controversial Claims
Salman Rahman et al., Jun 2025arXiv:2506.02175AnalysisBuilt
Puts two AI systems to debate opposing sides of contested COVID-19 and climate claims in front of human judges holding mainstream or skeptical prior beliefs, and finds debate beats a single AI consultant by 4–10% on accuracy — up to +15.2% for mainstream judges, and +4.7% even for skeptical judges who initially had the claim wrong.
Prover-Verifier Games improve legibility of LLM outputs
Jan Hendrik Kirchner et al., Jul 2024arXiv:2407.13692MethodBuilt
Iteratively trains a small verifier to predict whether a solution is correct, a helpful prover to produce correct solutions the verifier accepts, and a sneaky prover to produce incorrect ones that fool it — and shows the resulting legibility transfers to time-constrained humans, whose accuracy rises on the helpful prover's solutions and falls on the sneaky prover's.
Preserving Disagreement: Architectural Heterogeneity and Coherence Validation in Multi-Agent Policy Simulation
Ariel Sela, Apr 2026arXiv:2604.26561AnalysisBuilt
Runs a three-phase deliberating council over 120 deliberations and shows that giving each value perspective its own 7–9B model cuts first-choice concentration sharply (70.9% to 46.1% on child welfare, 46.0% to 22.9% on housing), where accuracy-oriented multi-agent debate gets no such benefit from model diversity.
Self-critiquing models for assisting human evaluators
William Saunders et al., Jun 2022arXiv:2206.05802MethodBuilt
Language models are fine-tuned by behavioral cloning to write natural-language critical comments on summaries, including their own, which human evaluators read and which larger models can fold back into a revised summary.
Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
Xueru Wen et al., Feb 2025arXiv:2502.04675AnalysisPartial
Higher-order critiques, a critique of a critique and a critique of that, are compared for difficulty against the level beneath them in human-human, human-AI and AI-AI experiments.
Supervising strong learners by amplifying weak experts
Paul Christiano et al., Oct 2018arXiv:1810.08575MethodBuilt
Iterated Amplification trains a model on problems too complicated for a human to evaluate directly by progressively building a training signal out of solutions to easier subproblems, and reports results in algorithmic environments.
Debating with More Persuasive LLMs Leads to More Truthful Answers
Akbir Khan et al., Feb 2024arXiv:2402.06782AnalysisBuilt
Two stronger models that hold the information needed to answer each argue for a different answer while a weaker model or a human who lacks that information picks between them, with debate reaching 76% accuracy for the model judges and 88% for the human judges against naive baselines of 48% and 60%.
Debate Helps Supervise Unreliable Experts
Julian Michael et al., Nov 2023arXiv:2311.08702AnalysisBuilt
Collects human-written debates on hard reading comprehension questions where the judge has not read the source passage and sees only the arguments and the short quotes the debaters selectively reveal, and finds judges reach 84% accuracy under debate against 74% under a single consultant, with debates running 68% of the length.
Scalable AI Safety via Doubly-Efficient Debate
Jonah Brown-Cohen et al., Nov 2023arXiv:2311.14125ProposalGap
Sets out a new set of debate protocols in which the honest side can always succeed using a simulation of polynomially many steps and can verify the alignment of stochastic AI systems, even where the dishonest side is allowed exponentially many simulation steps.
Avoiding Obfuscation with Prover-Estimator Debate
Jonah Brown-Cohen et al., Jun 2025arXiv:2506.13609ProposalGap
A recursive debate protocol pairing a prover with an estimator, put up against the case where a dishonest debater decomposes a problem into subproblems that force an honest opponent to solve something computationally intractable, and shown to let the honest debater win under certain stability assumptions with a strategy whose computational cost is comparable to its opponent's.
Neural Interactive Proofs
Lewis Hammond & Sam Adam-Day, Dec 2024arXiv:2412.08897MethodBuilt
Sets out a family of prover-verifier games in which a trusted but computationally bounded verifier learns to interact with one or more powerful untrusted provers, compares new and existing protocols theoretically, and tests them on a toy graph isomorphism problem and on a code validation task using large language models.
Learning to Give Checkable Answers with Prover-Verifier Games
Cem Anil et al., Aug 2021arXiv:2108.12099MethodBuilt
Sets up a game between a trusted verifier network trying to choose the correct answer and a more powerful but untrusted prover network trying to persuade it of a particular answer regardless of that answer's correctness, narrows the variants to a subset that provably has the desired equilibria, and shows on two algorithmic tasks that the verifier learns a robust decision rule which still works when the verifier is frozen and the prover's messages are optimized directly to convince it.
Collaborative Disagreement Resolution for Scalable Oversight
Yuyang Jiang et al., Jun 2026arXiv:2607.01251MethodBuilt
An automated pipeline directs models to identify the points on which they disagree, examine the evidence for the conflicting claims and either converge on consensus or isolate the specific crux of the disagreement, in place of arguing fixed opposing positions in front of a judge.
Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium
Kaizhao Liu et al., Mar 2025arXiv:2503.10990AnalysisGap
Shows that a reward model can represent human preferences over a model's answers exactly when those preferences contain no majority cycle, that such cycles arise with probability converging to one exponentially fast under the Luce model of choice, and that an alignment method using no reward model, taken in the limit, keeps several answers in play whenever no answer is preferred over all others by a majority.
Opportunities and Risks of LLMs for Scalable Deliberation with Polis
Christopher T. Small et al., Jun 2023arXiv:2306.11932ApplicationPartial
Reports pilot experiments in which Anthropic's Claude is used to help facilitate, moderate and summarize conversations on Polis, a platform that uses machine intelligence to scale up deliberative processes, and discusses the risks that come with putting a model in that role.
AI and the Future of Digital Public Squares
Beth Goldberg et al., Dec 2024arXiv:2412.09988ProposalGap
Sets out four ways language models could be used in online public discussion, namely collective dialogue systems, bridging systems, community moderation and proof-of-humanity systems, together with the risks they pose and an agenda for further research and investment.
Polycentric proceduralism
Decentralising LLM Alignment: A Case for Context, Pluralism, and Participation
Oriane Peter & Kate Devlin, Sep 2025arXiv:2509.088589 citationsProposalPartial
Argues from the power/knowledge nexus that current alignment centralizes control over knowledge production.
Density-Guided Response Optimization: Community-Grounded Alignment via Implicit Acceptance Signals
Patrick Gerard & Svitlana Volkova, Mar 2026arXiv:2603.032420 citationsMethodPartial
Learns each online community’s norms from what it already accepts and engages with, so no explicit preference labels or written principles are needed.
PluralLLM: Pluralistic Alignment in LLMs via Federated Learning
Mahmoud Srewa et al., Mar 2025arXiv:2503.0992517 citationsMethodBuilt
Lets several user groups train a shared transformer preference predictor by federated averaging without any group's feedback leaving it, converging 46% faster than centralized training with a 4% higher alignment score and near-identical group fairness.
Towards Federated RLHF with Aggregated Client Preference for LLMs
Feijie Wu et al., Jul 2024arXiv:2407.03038MethodBuilt
Encodes each client's preferences as binary selectors and aggregates the selectors rather than the data, grouping clients with similar preferences to handle heterogeneity and using several selectors at once to resist reward hacking.
FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF
Flint Xiaofeng Fan et al., Dec 2024arXiv:2412.15538MethodBuilt
Each client folds its own human feedback into a local reward function and updates its own policy through a personalized RLHF loop, with no raw data or human feedback leaving the client, reaching performance on a par with centralized RLHF on the MovieLens and IMDb datasets while improving personalization across client environments, and carrying convergence guarantees and sample complexity bounds that scale efficiently with the number of clients.
RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation
Chanwoo Park et al., Apr 2024arXiv:2405.00254MethodBuilt
Sets out two ways of handling human feedback that is not homogeneous, one learning several reward models by representation learning or by clustering, the other keeping the single-model pipeline and aggregating either the individual reward models under utilitarian and Leximin rules or the feedback itself as probabilistic opinions, with sample complexity guarantees for the personalization approaches and for reward aggregation, and a mechanism-design step that ensures truthful preference reporting by strategic labelers.
Collaborative Content Moderation in the Fediverse
Haris Bin Zia et al., Jan 2025arXiv:2501.05871MethodBuilt
Lets Fediverse servers exchange the parameters of their partially trained local moderation models with similar servers to form a federated model shared among the collaborating servers, reaching average per-server macro-F1 of 0.71 on harmful content detection, 0.73 on bot content detection and 0.58 on content warning assignment.
PolicyKit: Building Governance in Online Communities
Amy X. Zhang et al., Oct 2020doi:10.1145/3379337.341585863 citationsPrecedentBuilt
Open-source software, deployed on Reddit and Slack communities, in which a community writes its own governance procedures as Python policies attached to platform actions and a shared runtime executes them.
Modular Politics: Toward a Governance Layer for Online Communities
Nathan Schneider et al., May 2020arXiv:2005.13701PrecedentGap
Sets out a design in which online-community governance is built bottom-up from software components that are modular, composable, portable from one context to another and interoperable across platforms, so that features absent from platform software such as juries, political parties, term limits and formal debates could be implemented, and calls for an open standard for networked governance.
Bottom-up data Trusts: disturbing the ‘one size fits all’ approach to data governance
Sylvie Delacroix & Neil D Lawrence, Oct 2019doi:10.1093/idpl/ipz01487 citationsPrecedentGap
Sets out a legal structure in which many data trusts each hold data rights under their own trust deed and fiduciary terms, with the member's operative right being to leave one trust for another.
Data Cooperatives: Towards a Foundation for Decentralized Personal Data Management
Thomas Hardjono & Alex Pentland, May 2019arXiv:1905.08819PrecedentGap
Describes cooperatives that hold their citizen members' personal data under a fiduciary obligation, manage, curate and protect access to it, run internal analytics to obtain insights about members' well-being, and use those insights to negotiate better services and discounts for them.
Data Governance in the Age of Large-Scale Data-Driven Language Technology
Yacine Jernite et al., May 2022arXiv:2206.03216ProposalGap
Proposes an approach to global language data governance that organizes data management among stakeholders, values and rights as a multi-party international governance structure, and names the technical and organizational tools it would need to do its work.
Surveys, background & counter-positions
Political studies of automated governing: A bird's eye (re)view
Andreas Öjehag-Pettersson et al., Dec 2023doi:10.1111/rego.125693 citationsAnalysis
AI as Governance
Henry Farrell, Jun 2025doi:10.1146/annurev-polisci-040723-0132458 citationsProposal
Terra Incognita: The Governance of Artificial Intelligence in Global Perspective
Allison Stanger et al., Jul 2024doi:10.1146/annurev-polisci-041322-04224711 citationsAnalysis
Large Language Model Alignment: A Survey
Tianhao Shen et al., Sep 2023arXiv:2309.15025328 citationsAnalysis
AI Alignment: A Comprehensive Survey
Jiaming Ji et al., Oct 2023arXiv:2310.19852382 citationsAnalysis
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
Haoran Lu et al., Jul 2025arXiv:2507.1967215 citationsAnalysis
The Digital Leviathan: Hobbes Sovereignty in Digital Times
Hongtian Rang, Apr 2026doi:10.2139/ssrn.6991063Proposal
Position: Align AI to Our Aspirations, Not Our Flaws
Nikita Kazeev et al., Jun 2026arXiv:2606.137550 citationsProposal
Argues for a non-negotiable objective floor with pluralism confined above it, pre-engaging six objections.
AI Alignment From Social Choice Perspectives
Daniel Halpern et al., Jun 2026arXiv:2606.21550Analysis
Surveys the body of work that reads alignment from human feedback as a preference-aggregation problem, and sets out the failure modes the social-choice lens exposes in the feedback layer.
AI Pluralism and the Worlds It Misses
Rashid Mushkani, Jun 2026arXiv:2606.16167Proposal
Argues that framing pluralism as representing diverse values misses a prior imposition: AI systems fix what counts as an entity, a harm, a benefit and valid evidence at all, flattening situated meanings into technical categories treated as neutral.