Papers

239 of 239 papers — 93 are methods, the rest benchmarks, applications, resources, analyses, proposals and precedents

Constitutionalism

Claude's Constitution

Amanda Askell et al. (Anthropic), Jan 2026

anthropic.comResourceBuilt

Deliberative Alignment: Reasoning Enables Safer Language Models

Melody Y. Guan et al. (OpenAI), Dec 2024

arXiv:2412.16339277 citationsMethodPartial

Teaches the safety spec directly and trains the model to recall and reason over it before answering.

Constitutional AI: Harmlessness from AI Feedback

Yuntao Bai et al. (Anthropic), Dec 2022

arXiv:2212.080733,425 citationsMethodBuilt

Self-critique and revision against a written list of principles, followed by RLAIF.

Specific versus General Principles for Constitutional AI

Sandipan Kundu et al. (Anthropic), Oct 2023

arXiv:2310.1379855 citationsMethodBuilt

Tests whether a single general principle can substitute for a list of specific rules.

Model Spec

OpenAI, May 2024

model-spec.openai.comResourceBuilt

Rule-Based Rewards for Language Model Safety

Tong Mu et al., Nov 2024

arXiv:2411.01111MethodBuilt

C3AI: Crafting and Evaluating Constitutions for Constitutional AI

Yara Kyrychenko et al., Apr 2025

doi:10.1145/3696410.37147058 citationsMethodBuilt

Selects and structures the principles that go into a constitution before fine-tuning, then checks whether the fine-tuned model actually follows them — finding that positively framed, behavior-based principles match human preferences best while the trained models follow negatively framed ones better.

How Well Do Models Follow Their Constitutions?

Arya Jakkli et al., May 2026

arXiv:2605.24229BenchmarkBuilt

Decomposes each lab's published specification into atomic testable tenets — 205 for Anthropic's constitution, 197 for OpenAI's Model Spec — generates multi-turn adversarial scenarios against them, and reports violation rates falling generation over generation: 15.0% to 2.0% across the Claude family, 11.7% to 3.6% across the GPT family, with the severity ceiling dropping from 10/10 to 7/10.

EigenBench: A Comparative Behavioral Measure of Value Alignment

Jonathn Chang et al., Sep 2025

arXiv:2509.01938BenchmarkBuilt

Scores a group of models against one written constitution by having each model judge the others' outputs across many scenarios and aggregating those judgments with EigenTrust, returning a vector of alignment scores rather than a single pass mark — with no ground-truth labels, because the traits it measures are ones reasonable judges disagree about.

Policy-as-Prompt: Turning AI Governance Rules into Guardrails for AI Agents

Gauri Kholkar & Ratinder Ahuja, Sep 2025

arXiv:2509.23994MethodBuilt

Reads an organization's product requirement, design and code documents, builds a policy tree whose every node links back to the source text it came from, and compiles that tree into prompt-based classifiers that monitor an agent at run time.

RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Harrison Lee et al., Sep 2023

arXiv:2309.00267MethodBuilt

Trains the reward model on preference labels produced by an off-the-shelf language model instead of by people, across summarization, helpful dialogue and harmless dialogue.

Principle-Driven Self-Alignment of Language Models from Scratch with Minimal Human Supervision

Zhiqing Sun et al., May 2023

arXiv:2305.03047MethodBuilt

Prompts a base model with sixteen written principles and five worked examples so that it generates its own compliant answers, then fine-tunes on those answers until the principles are no longer needed at inference.

Improving alignment of dialogue agents via targeted human judgements

Amelia Glaese et al., Sep 2022

arXiv:2209.14375MethodBuilt

Breaks the requirements for good dialogue into natural-language rules, asks human raters to judge each rule separately, and trains rule-conditional reward models on those per-rule judgments.

Can LLMs Follow Simple Rules?

Norman Mu et al., Nov 2023

arXiv:2311.04235BenchmarkBuilt

Puts a model through 14 simple text scenarios in which it is instructed to obey various rules while interacting with a user, and uses a program rather than a human rater to determine whether any rule was broken in the conversation, finding that almost all current proprietary and open models struggle even on straightforward test cases and that simple optimization attacks significantly raise failure rates.

SpecEval: Evaluating Model Adherence to Behavior Specifications

Ahmed Ahmed et al., Sep 2025

arXiv:2509.02464BenchmarkBuilt

Parses a provider's published specification into individual behavioral statements, generates prompts targeted at each one, and has the provider's own models judge whether its outputs comply, run over 16 models from six developers across more than 100 behavioral statements and finding compliance gaps of up to 20 percent.

Stress-Testing Model Specs Reveals Character Differences among Language Models

Jifan Zhang et al., Oct 2025

arXiv:2510.07686BenchmarkBuilt

Generates scenarios that force a choice between pairs of legitimate principles that cannot both be satisfied, evaluates twelve frontier models from Anthropic, OpenAI, Google and xAI on them, and identifies over 70,000 cases of significant behavioral divergence, divergence that strongly predicts underlying problems in the specifications themselves, with the qualitative analysis naming direct contradictions and interpretive ambiguities among them.

Does Claude's Constitution Have a Culture?

Parham Pourdavood, Mar 2026

arXiv:2603.28123AnalysisBuilt

Puts Claude Sonnet through 55 World Values Survey items selected for high cross-cultural variance across six value domains, administered both as direct survey questions and as naturalistic advice-seeking scenarios, and compares the answers with country-level data from 90 nations.

Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming

Mrinank Sharma et al., Jan 2025

arXiv:2501.18837MethodBuilt

Trains classifier safeguards on synthetic data generated by prompting language models with natural-language rules, a constitution, specifying permitted and restricted content, and reports that in over 3,000 estimated hours of red teaming no red teamer found a universal jailbreak that could extract information from an early classifier-guarded model at a level of detail similar to an unguarded model across most target queries, at a cost of an absolute 0.38% increase in production-traffic refusals and 23.7% inference overhead.

Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations

Hakan Inan et al., Dec 2023

arXiv:2312.06674MethodBuilt

A Llama2-7b model instruction-tuned on a small hand-gathered dataset to classify both prompts and responses against a written safety risk taxonomy, released with open weights, matching or exceeding available content moderation tools on the OpenAI Moderation Evaluation dataset and ToxicChat, and able to take a different taxonomy at the input for zero-shot or few-shot use.

SafetyAnalyst: Interpretable, Transparent, and Steerable Safety Moderation for AI Behavior

Jing-Jing Li et al., Oct 2024

arXiv:2410.16665MethodBuilt

Uses chain-of-thought reasoning to analyze a candidate AI behavior into a structured harm-benefit tree of the harmful and beneficial actions and effects it may lead to, each labeled for likelihood, severity and immediacy of impact on stakeholders, then aggregates them into a harmfulness score through 28 fully interpretable weight parameters, in an open-source prompt safety classifier distilled from 18.5 million harm-benefit features generated by frontier models on 19k prompts that reaches average F1 0.81 where existing moderation systems score below 0.72.

The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Eric Wallace et al., Apr 2024

arXiv:2404.13208MethodBuilt

Defines an explicit priority ordering over the instructions a model receives, so that a system prompt from an application developer outranks text from untrusted users and third parties, and generates training data teaching the model to ignore lower-privileged instructions selectively, which applied to GPT-3.5 sharply increases robustness including to attack types not seen during training, with minimal degradation on standard capabilities.

SPRI: Aligning Large Language Models with Context-Situated Principles

Hongli Zhan et al., Feb 2025

arXiv:2502.03397MethodBuilt

Generates guiding principles in real time for each individual input query with minimal or no human effort and uses them to align the response, deriving principles in a complex domain-specific task that perform on a par with expert-crafted ones, turning them into instance-specific rubrics that outperform prior LLM-as-a-judge frameworks, and using them to generate synthetic supervised fine-tuning data that substantially improves truthfulness.

Contestatory constitutionalism

Public Constitutional AI

Gilad Abiri, 2024

arXiv:2406.166969 citationsProposalGap

Argues AI governance needs democratic legitimation, and proposes a constitution plus an accumulated public case law.

Cognitive Comparability and the Limits of Governance: Evaluating Authority Under Radical Capability Asymmetry

Tony Rost, Apr 2026

arXiv:2604.027201 citationProposalPartial

Sets out a six-dimension test — legitimacy, accountability, corrigibility, non-domination, subsidiarity, institutional resilience — first on existing non-majoritarian institutions and then on a hypothetical bounded superintelligent authority.

Infrastructuring Contestability: A Framework for Community-Defined AI Value Pluralism

Andreas Mayer, Jul 2025

arXiv:2507.051870 citationsProposalPartial

Proposes Community-Defined AI Value Pluralism: communities author their own value profiles, and the system exposes the interfaces through which those profiles can be challenged and revised.

Deference by Design: Pluralistic Alignment Is an Interface Problem

Jul 2026

pluralistic-alignment.github.ioProposalPartial

Statutory Construction and Interpretation for Artificial Intelligence

Luxi He et al., Sep 2025

arXiv:2509.01186MethodBuilt

Shows that the same natural-language principle admits several defensible readings that push model behavior apart, then builds two mechanisms borrowed from law — a pipeline that revises ambiguous rules, the way an agency reworks a regulation, and prompt-based interpretive constraints that play the role legal canons play in guiding discretion — and shows on 5,000 WildChat scenarios that both raise agreement across a panel of reasonable interpreters.

Alignment as Jurisprudence

Nicholas Caputo, May 2026

arXiv:2605.08416ProposalGap

Reads the alignment problem as the jurisprudential one: both fields try to fix in language how a powerful decider will act in situations nobody has yet seen, and both split over whether to bind the decider to rules or to accumulated cases.

AppealMod: Inducing Friction to Reduce Moderator Workload of Handling User Appeals

Shubham Atreja et al., Apr 2024

doi:10.1145/36372967 citationsPrecedentBuilt

An appeals system for a Reddit community of millions of subscribers in which a user appealing a ban must first supply further information through a bot, which then hands the exchange to human moderators as a structured record.

Contestable AI by Design: Towards a Framework

Kars Alfrink et al., Aug 2022

doi:10.1007/s11023-022-09611-z82 citationsProposalGap

A framework of design features and organizational practices, drawn from a systematic review, for building AI systems that the people subject to them can dispute.

Contestable Camera Cars: A Speculative Design Exploration of Public AI That Is Open and Responsive to Dispute

Kars Alfrink et al., Apr 2023

doi:10.1145/3544548.358098437 citationsApplicationPartial

Works the contestable-AI framework into a concept for a municipal camera car, with notice of the automated assessment, affordances for disputing it, appeal routing to civil servants and an audit trail, rendered as a concept video and evaluated through interviews with 17 civil servants operating AI in a large European city.

ConGaIT: A Clinician-Centered Dashboard for Contestable AI in Parkinson's Disease Care

Phuc Truong Loc Nguyen & Thanh Hung Do, Jul 2025

arXiv:2507.22300ApplicationBuilt

A dashboard for Parkinson's disease gait analysis in which a clinician registers structured disagreement with the model's reading through a Contest and Justify interaction, backed by visual explanations, role-based feedback and traceable justification logs.

Explainable AI Systems Must Be Contestable: Here's How to Make It Happen

Catarina Moreira et al., Jun 2025

arXiv:2506.01662BenchmarkPartial

Gives a formal definition of contestability in explainable AI, a modular set of by-design and post-hoc mechanisms spanning human-centered interfaces, technical architectures, legal processes and organizational workflows, and the Contestability Assessment Scale, a composite metric built on more than twenty quantitative criteria.

FLARE-AI: Flaw Reporting for AI

Shayne Longpre et al., Jun 2026

arXiv:2606.31567ResourceBuilt

An open-source reporting system that takes a single submission about a flaw in a deployed AI system, asks follow-up questions conditioned on what the reporter has already said, and can then send a standardized machine-readable report from that one submission to several developers, coordinators and incident registries at once.

In-House Evaluation Is Not Enough: Towards Robust Third-Party Flaw Disclosure for General-Purpose AI

Shayne Longpre et al., Mar 2025

arXiv:2503.16861ProposalGap

A collaboration of software security, machine learning, law, social science and policy researchers sets out three things general-purpose AI lacks: standard flaw report formats with rules of engagement, broadly scoped disclosure programs borrowed from bug bounties with legal safe harbours for reporters, and infrastructure to distribute a report across the stakeholders it concerns.

Designing Incident Reporting Systems for Harms from General-Purpose AI

Kevin Wei & Lennart Heim, Mar 2026

doi:10.1609/aaai.v40i44.411390 citationsProposalGap

A seven-part design framework for AI incident reporting schemes, covering the policy goal, the actors who submit and receive, the type of incident, how far the risk materialized, enforcement, reporter anonymity and what happens after a report, drawn from nine case studies of incident reporting in safety-critical industries and worked into design specifications for a United States regime for general-purpose AI.

Preventing Repeated Real World AI Failures by Cataloging Incidents: The AI Incident Database

Sean McGregor, Nov 2020

arXiv:2011.08512ResourceBuilt

A database of real-world AI failures, started by an industrial and non-profit cooperative, with faceted and full-text search over more than 1,000 archived incident reports.

Lessons for Editors of AI Incidents from the AI Incident Database

Kevin Paeth et al., Sep 2024

arXiv:2409.16425AnalysisPartial

A review of over 750 incidents in the AI Incident Database and of two independent taxonomies applied to them, reporting the patterns that make incidents hard to index and the mitigations editors can use when cause, extent of harm, severity or the technical details of the systems involved are uncertain.

Audit Trails for Accountability in Large Language Models

Victor Ojewale et al., Jan 2026

arXiv:2601.20727MethodBuilt

An append-only, tamper-evident ledger for language model deployments that records lifecycle events such as models, data, training and evaluation runs, deployments and monitoring, alongside the approvals, waivers and attestations that authorized them, released as an open-source Python implementation that emits these records from existing workflows.

The DSA Transparency Database: Auditing Self-reported Moderation Actions by Social Media

Amaury Trujillo et al., May 2025

doi:10.1145/371108515 citationsPrecedentBuilt

An audit of all 353.12 million statements of reasons that the eight largest social platforms in the EU submitted to the Digital Services Act Transparency Database in its first hundred days, comparing across platforms the grounds given for each decision, the types of restriction imposed, the timeliness of the actions and the use of automation.

Contestable AI needs Computational Argumentation

Francesco Leofante et al., May 2024

arXiv:2405.10729ProposalGap

A position paper arguing that a contestable AI system must be able to hold an exchange with humans or other machines, explaining its output and its reasoning step by step, weighing the grounds offered against it, and revising how it decides when a challenge succeeds, and that computational argumentation is the technique suited to supporting this.

Contestability in Quantitative Argumentation

Xiang Yin et al., Jul 2025

arXiv:2507.11323MethodBuilt

A method that takes an argument network with weighted supporting and attacking links, computes how sensitive a chosen conclusion's strength is to each individual link weight, and then adjusts the weights step by step until that conclusion reaches a desired strength.

Consociational codification

Legal Alignment for Safe and Ethical AI

Noam Kolt et al., Jan 2026

arXiv:2601.041758 citationsProposalPartial

Surveys legal alignment with a taxonomy of three pathways: complying with rules, adapting legal interpretation, and using legal concepts as blueprints.

International Governance of Civilian AI: A Jurisdictional Certification Approach

Robert Trager et al., Aug 2023

arXiv:2308.15514ProposalGap

Proposes an International AI Organization that certifies national jurisdictions — not firms and not individual models — against shared oversight standards, with certified states barring imports of goods whose supply chains embody AI from uncertified jurisdictions and controlling exports of AI inputs such as specialized hardware to them.

SafeWorld: Geo-Diverse Safety Alignment

Da Yin et al., Dec 2024

arXiv:2412.06483MethodBuilt

Builds a 2,342-query benchmark grounded in human-verified cultural norms and legal policies from 50 countries and 493 regions or ethnic groups, then trains SafeWorldLM by direct preference optimization to answer the same question differently by context and to cite the norm or policy it is applying.

The Law-Following AI Framework: Legal Foundations and Technical Constraints. Legal Analogues for AI Actorship and technical feasibility of Law Alignment

Katalina Hernandez Delgado, Sep 2025

arXiv:2509.08009AnalysisGap

Tests the Law-Following AI proposal against existing legal categories, showing that the law already recognizes actors who bear duties without full personhood, and then asks whether law alignment is technically feasible as a superordinate objective.

Law-Following AI: Designing AI Agents to Obey Human Laws

Cullen O'Keefe et al., 2025

doi:10.2139/ssrn.52426436 citationsProposalGap

A design specification for AI agents in which compliance with a designated body of positive law is a superordinate objective that no other goal may override, backed by a legal construct of AI actorship without legal personhood and by duty-bearing and liability-channelling mechanisms to enforce it.

Regulatory Markets for AI Safety

Jack Clark & Gillian K. Hadfield, Dec 2019

arXiv:2001.00078ProposalGap

Proposes global regulatory markets as a model for achieving AI safety, sketches the model in general terms with an overview of its costs and benefits, and works it through on one risk: adversarial attacks on AI models employed in commercial drones.

ML-Bench&Guard: Policy-Grounded Multilingual Safety Benchmark and Guardrail for Large Language Models

Yunhan Zhao et al., May 2026

arXiv:2605.00689MethodBuilt

Derives risk categories and fine-grained rules from jurisdiction-specific legal texts and uses them to generate a safety benchmark covering 14 languages, then builds on that benchmark a diffusion-model guardrail in two sizes, a 1.5B one for fast safe or unsafe checks and a 7B one that assesses compliance against a policy supplied to it and explains its verdict, reported as consistently outperforming 11 guardrail baselines across six existing multilingual benchmarks and its own.

SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia

Panuthep Tasawong et al., Feb 2026

arXiv:2602.01618MethodBuilt

Generates region-specific safety data for Southeast Asia with an agentic data-generation pipeline rather than by machine-translating English datasets, and trains a family of safeguard models on it that detect regionally sensitive or harmful content better than existing safeguards while maintaining strong general safety performance.

PluRule: A Benchmark for Moderating Pluralistic Communities on Social Media

Zoher Kachwala et al., May 2026

arXiv:2605.17187BenchmarkBuilt

Poses moderation as a multiple-choice task in which a model is given a comment and its surrounding context and must identify which specific rule, if any, it violates, across 13,371 rule violations from 1,989 Reddit communities spanning 2,885 rules in nine languages, and finds that even GPT-5.2 with high reasoning performs only slightly better than a trivial baseline.

COMPL-AI Framework: A Technical Interpretation and LLM Benchmarking Suite for the EU Artificial Intelligence Act

Philipp Guldimann et al., Oct 2024

arXiv:2410.07959BenchmarkBuilt

A translation of the EU AI Act's broad regulatory requirements into measurable technical requirements for language models, together with an open-source Act-centered benchmarking suite implementing them, run over 12 prominent models.

The Sandbox Configurator: A Framework to Support Technical Assessment in AI Regulatory Sandboxes

Alessio Buscemi et al., Sep 2025

arXiv:2509.25256ProposalPartial

A modular open-source framework in which a user selects domain-relevant tests from a shared library and generates a customized sandbox environment with integrated dashboards, for AI systems assessed under the oversight of a Competent Authority.

Navigating Global AI Regulation: A Multi-Jurisdictional Retrieval-Augmented Generation System

Courtney Ford et al., Apr 2026

arXiv:2604.25448MethodBuilt

A retrieval-augmented system over 242 regulatory documents from 68 jurisdictions that chunks each document according to its type to preserve legal structure, routes a query using detected entities and citation metadata, and ranks enacted legislation above policy and secondary sources before answering.

GoldCoin: Grounding Large Language Models in Privacy Laws via Contextual Integrity Theory

Wei Fan et al., Jun 2024

arXiv:2406.11149MethodBuilt

A framework that generates synthetic scenarios grounded in privacy statutes, using contextual integrity as the bridge between statute and situation, so that a language model given them recognizes privacy violations in real court cases.

Authenticated Delegation and Authorized AI Agents

Tobin South et al., Jan 2025

arXiv:2501.09674ProposalGap

It sets out an extension of OAuth 2.0 and OpenID Connect with agent-specific credentials and metadata that record which human or organization an AI agent acts on behalf of and what permissions it has been given, together with a scheme for turning permissions written in natural language into auditable access-control configurations.

IDs for AI Systems

Alan Chan et al., Jun 2024

arXiv:2406.12137ProposalGap

It proposes giving identifiers to individual instances of AI systems, such as a particular chat session, with associated information made accessible to parties seeking to interact with that instance.

Infrastructure for AI Agents

Alan Chan et al., Jan 2025

arXiv:2501.10114ProposalGap

It sets out the idea of agent infrastructure, meaning technical systems and shared protocols external to agents that attribute actions to particular agents or people, shape how agents interact, and detect and remedy harmful actions, and catalogs research directions for each function.

Liability, Ethics, and Culture-Aware Behavior Specification using Rulebooks

Andrea Censi et al., Feb 2019

arXiv:1902.09355MethodBuilt

It defines a rulebook as a pre-ordered set of rules, each akin to a violation metric over possible outcomes, whose priority ordering imposes a pre-order on those outcomes, and derives which operations on rulebooks preserve constraints introduced earlier.

Behavioral Use Licensing for Responsible AI

Danish Contractor et al., Jun 2022

doi:10.1145/3531146.353314347 citationsProposalBuilt

It sets out licences that attach enumerated use restrictions to model weights, so the terms on which a released model may be used travel with the artifact.

SILO Language Models: Isolating Legal Risk In a Nonparametric Datastore

Sewon Min et al., Aug 2023

arXiv:2308.04430MethodBuilt

It trains a language model only on public domain and permissively licensed text, 228 billion tokens of it, and augments it with a separate nonparametric datastore of higher-risk material queried only at inference, so a generation can be attributed to a sentence in the store and content can be removed from the store without retraining.

Precedential adjudication

Case Law Grounding: Using Precedents to Align Decision-Making for Humans and AI

Quan Ze Chen & Amy X. Zhang, Oct 2023

arXiv:2310.070199 citationsMethodPartial

Human-led and LLM-prompted versions of case-law grounding, evaluated on content moderation and toxicity rating.

Case Repositories: Towards Case-Based Reasoning for AI Alignment

K. J. Kevin Feng et al., Nov 2023

arXiv:2311.1093417 citationsProposalPartial

Four-step pipeline for assembling a case repository: seed cases, expert-elicited dimensions, LLM-generated variations, public judgment.

Inverse Constitutional AI: Compressing Preferences into Principles

Arduin Findeis et al., Jun 2024

arXiv:2406.0656044 citationsMethodBuilt

Derives an interpretable principle set from a preference dataset: constitutional AI run backwards.

Reasoning over Precedents Alongside Statutes: Case-Augmented Deliberative Alignment for LLM Safety

Can Jin et al., Jan 2026

arXiv:2601.080003 citationsMethodBuilt

Compares spelling safety rules out at length against demonstrating them through decided cases, finds that extensive codes improve harmlessness only inconsistently while systematically degrading helpfulness, and then trains a model by reinforcement learning on its own case-augmented safety reasoning chains.

CARO: Chain-of-Analogy Reasoning Optimization for Robust Content Moderation

Bingzhe Wu et al., Apr 2026

arXiv:2604.10504MethodBuilt

Trains a moderation model in two stages to reason by explicit analogy to retrieved past cases — first bootstrapping analogy chains by retrieval, then optimizing them — so it stops relying on the surface shortcuts that mislead it on ambiguous content.

IterAlign: Iterative Constitutional Alignment of Large Language Models

Xiusi Chen et al., Mar 2024

arXiv:2403.18341MethodBuilt

Red-teams a model, reads the failures, writes new constitutional principles that would have prevented them, and aligns the model to those — so the constitution is discovered from the model's own failures rather than written in advance.

Decoding Human Preferences in Alignment: An Improved Approach to Inverse Constitutional AI

Carl-Leander Henneking & Claas Beger, Jan 2025

arXiv:2501.17112MethodBuilt

Improves the Inverse Constitutional AI algorithm — how candidate principles are generated, clustered and embedded — so the constitution recovered from a preference dataset is a more faithful account of what the raters were actually doing.

Extensionally defining principles and cases in ethics: An AI model

Bruce M. McLaren, Nov 2003

doi:10.1016/S0004-3702(03)00135-868 citationsMethodBuilt

SIROCCO takes a new professional-ethics case and returns the past decisions and the principles that bear on it, working in two stages over 500 cases decided by the National Society of Professional Engineers' Board of Ethical Review.

Can Machines Learn Morality? The Delphi Experiment

Liwei Jiang et al., Oct 2021

arXiv:2110.07574MethodBuilt

Delphi is a neural model trained to make descriptive ethical judgments about everyday situations, so that "helping a friend" comes back good and "helping a friend spread fake news" does not.

Divergent precedent

Rules, Cases, and Reasoning: Positivist Legal Theory as a Framework for Pluralistic AI Alignment

Nicholas A. Caputo, Oct 2024

arXiv:2410.172713 citationsProposalGap

Dealing with Disagreements: Looking Beyond the Majority Vote in Subjective Annotations

Aida Mostafazadeh Davani et al., Oct 2021

arXiv:2110.05719MethodBuilt

It trains one model with a separate prediction subtask for each annotator over a shared learned representation, so the model predicts what each individual annotator would say instead of a majority label.

DICES Dataset: Diversity in Conversational AI Evaluation for Safety

Lora Aroyo et al., Jun 2023

arXiv:2306.11247BenchmarkBuilt

It is a safety-rating dataset for conversational AI in which each item is rated many times over, fine-grained demographic information about raters is recorded, and votes are encoded as distributions across demographics rather than reduced to one label.

LeWiDi-2025 at NLPerspectives: Third Edition of the Learning with Disagreements Shared Task

Elisa Leonardelli et al., Oct 2025

arXiv:2510.08460BenchmarkBuilt

It runs a shared task in which systems are scored on how well they reproduce human disagreement across four datasets covering paraphrase identification, irony detection, sarcasm detection and natural language inference, under two paradigms: predicting population-level distributions of judgments, and predicting the interpretations of individual annotators.

Diverging Preferences: When do Annotators Disagree and do Models Know?

Michael JQ Zhang et al., Oct 2024

arXiv:2410.14632MethodBuilt

It sorts the reasons annotators of preference data disagree into ten categories across four high-level classes, and shows that standard Bradley-Terry reward modeling and LLM-as-judge evaluation fail to account for divergence between annotators.

Reasoning with cases and hypotheticals in HYPO

Kevin D. Ashley, Jun 1991

doi:10.1016/0020-7373(91)90011-U115 citationsMethodBuilt

An implemented case-based reasoning system that retrieves relevant past cases through a claim lattice and produces a three-ply argument: precedents cited for one side, distinguished and counter-cited for the other, with manufactured hypotheticals used to probe a position.

Do LLMs Truly Understand When a Precedent Is Overruled?

Li Zhang et al., Oct 2025

arXiv:2510.20941BenchmarkBuilt

A benchmark of 236 U.S. Supreme Court case pairs on which state-of-the-art language models are asked to identify whether one case overrules the other, with performance broken out by era and by task format.

PolicyCraft: Supporting Collaborative and Participatory Policy Design through Case-Grounded Deliberation

Tzu-Sheng Kuo et al., Apr 2025

doi:10.1145/3706598.371386514 citationsPrecedentBuilt

A deployed web system in which community members submit concrete cases, deliberate over them and grow a community-owned policy document that is revised case by case, with the system tracking which clauses each case grounds and surfacing cases that contradict the policy as it stands.

Botender: Supporting Communities in Collaboratively Designing AI Agents through Case-Based Provocations

Tzu-Sheng Kuo et al., Sep 2025

arXiv:2509.25492MethodBuilt

A no-code system in which community members propose, iterate on and deploy the behavior of a language-model bot, using generated interaction scenarios as provocations to prompt discussion about what the bot should do.

Judgment Sieve: Reducing Uncertainty in Group Judgments through Interventions Targeting Ambiguity versus Disagreement

Quan Ze Chen & Amy X. Zhang, Sep 2023

doi:10.1145/36100749 citationsMethodBuilt

A measurement framework and experimental pipeline that splits the uncertainty in a group's judgments on content-moderation cases into ambiguity, which clarifying the case can reduce, and genuine disagreement, which it cannot, and applies a different intervention to each.

Parallel lineages

MoMoE: Mixture of Moderation Experts Framework for AI-Assisted Online Governance

Agam Goyal et al., May 2025

arXiv:2505.14483PrecedentBuilt

Runs seven community-specialized moderation experts plus five norm-violation experts under an allocator that picks which expert judges a given post, an aggregator, and an explainer that gives the reason in that community's terms.

Decoding the Rule Book: Extracting Hidden Moderation Criteria from Reddit Communities

Youngwoo Kim et al., Sep 2025

arXiv:2509.02926PrecedentBuilt

Recovers each subreddit's unwritten moderation standard from its own history of removals, as an interpretable score table of lexical expressions, and shows the extracted tables match neural moderators' performance while making the criteria comparable across communities.

Crossmod: A Cross-Community Learning-based System to Assist Reddit Moderators

Eshwar Chandrasekharan et al., Nov 2019

doi:10.1145/3359276132 citationsPrecedentBuilt

A moderation bot, released as open source and run live in a large subreddit, that learns from millions of moderator removal decisions taken across 100 communities and refers its judgments to the host community's own moderators for review, with a separate threshold set for each community.

The Internet's Hidden Rules

Eshwar Chandrasekharan et al., Nov 2018

doi:10.1145/3274301201 citationsPrecedentBuilt

Takes 2.8 million comments removed from 100 Reddit communities and induces from the removals alone which norms each community actually enforces, sorting them into platform-wide, cluster-level and single-community strata.

SPICA: Retrieving Scenarios for Pluralistic In-Context Alignment

Quan Ze Chen et al., Nov 2024

arXiv:2411.10912MethodBuilt

Retrieves few-shot examples for a model from a bank of past scenarios using metrics that weigh how groups differ rather than similarity alone; on an alignment task drawing inputs from four demographic groups (n = 544) the retrieved examples matched observed preferences more closely, and in an end-to-end evaluation (n = 120) it was rated above similarity-based retrieval, with groups gaining up to 0.16 points on a five-point scale and every group benefiting rather than only some.

Customize Multi-modal RAI Guardrails with Precedent-based predictions

Cheng-Fu Yang et al., Jul 2025

arXiv:2507.20503MethodBuilt

Judges whether an image breaches a user-defined content policy by conditioning on precedents, the recorded reasoning from earlier similar inputs collected by a critique-and-revise mechanism, rather than on the policy text, and reports better results than previous methods in both few-shot and full-dataset settings and better generalization to policies never seen in training.

Jury Learning: Integrating Dissenting Voices into Machine Learning Models

Mitchell L. Gordon et al., Apr 2022

doi:10.1145/3491102.3502004101 citationsMethodBuilt

A deep learning architecture that models each individual annotator conditioned on their group identity, paired with an interactive system in which a practitioner declares a jury — which groups, in what proportion — and reads off that jury's verdict along with the distribution of dissent inside it.

Convention equilibrium

Legible Normativity for AI Alignment: The Value of Silly Rules

Dylan Hadfield-Menell et al., Nov 2018

arXiv:1811.0126723 citationsPrecedentPartial

Shows that arbitrary but highly legible rules help agents develop the general capacity to recognize and follow norms.

Emergent social conventions and collective bias in LLM populations

Ariel Flint Ashery et al., Oct 2024

arXiv:2410.08948133 citationsAnalysisBuilt

A population of LLM agents playing a naming game spontaneously converges on shared conventions, with committed-minority tipping points.

Artificial Leviathan: Exploring Social Evolution of LLM Agents Through the Lens of Hobbesian Social Contract Theory

Gordon Dai et al., Jun 2024

arXiv:2406.1437311 citationsAnalysisBuilt

Cultural Evolution of Cooperation among LLM Agents

Aron Vallinder & Edward Hughes, Dec 2024

arXiv:2412.10270AnalysisBuilt

Plays a classic iterated Donor Game across generations of LLM agents that can observe their peers' recent behavior, and finds indirect reciprocity evolving very differently by base model — Claude 3.5 Sonnet societies reach substantially higher average scores than Gemini 1.5 ones.

Cultural evolution in populations of Large Language Models

Jérémy Perez et al., Mar 2024

arXiv:2403.08882ResourceBuilt

Provides an open-source framework for simulating cultural evolution in populations of LLM agents, letting the variables cultural evolution cares about — network structure, agent personality, and how social information is aggregated and transformed — be manipulated directly.

Evolution of Social Norms in LLM Agents using Natural Language

Ilya Horiguchi et al., Sep 2024

arXiv:2409.00993AnalysisBuilt

Rebuilds Axelrod's metanorm games with LLM agents that talk to each other in natural language, and shows the agents form and then enforce normative strategies through the dialogue itself.

Normative Modules: A Generative Agent Architecture for Learning Norms that Supports Multi-Agent Cooperation

Atrisha Sarkar et al., May 2024

arXiv:2405.19328MethodBuilt

Equips a generative agent with a normative module that learns, through interaction with peers, which of several candidate institutions a group treats as authoritative — then shows the agent can disregard non-authoritative ones, pick the authoritative one out of several, and reach more stable cooperation than agents without the module.

Emergence of Social Norms in Generative Agent Societies: Principles and Architecture

Siyue Ren et al., Mar 2024

arXiv:2403.08251MethodBuilt

An architecture in four modules that has agents in the Smallville sandbox create norms, hold them in an explicit representation, pass them on through conversation and observation, check them, and act on them in planning.

Cooperate or Collapse: Emergence of Sustainable Cooperation in a Society of LLM Agents

Giorgio Piatti et al., Apr 2024

arXiv:2404.16698BenchmarkBuilt

A simulation platform in which a society of LLM agents must balance drawing on a shared resource against sustaining it for future use, and in which the highest survival rate across the models tested is below 54%.

Project Sid: Many-agent simulations toward AI civilization

Altera. AL et al., Oct 2024

arXiv:2411.00114MethodBuilt

Runs 10 to over 1,000 AI agents together in Minecraft under an architecture that keeps their several output streams coherent in real time, and reports agents taking up specialized roles, adhering to and changing collective rules, and transmitting culture and religion.

Spurious normativity enhances learning of compliance and enforcement behavior in artificial agents

Raphael Köster et al., Jan 2022

doi:10.1073/pnas.210602811832 citationsAnalysisBuilt

A multi-agent reinforcement learning environment in which a taboo carrying no intrinsic cost is added to the norm set, with compliance, third-party punishment and enforcement skill measured across the trained populations.

A learning agent that acquires social norms from public sanctions in decentralized multi-agent settings

Eugene Vinitsky et al., Jun 2021

arXiv:2106.09012MethodBuilt

An agent architecture combining a classifier that sorts observed behavior into approved or disapproved with a motivation to punish in accord with the group, trained in a regime where every agent can see all sanctioning events but learning is otherwise decentralized.

Contested custom

Pluralistic Alignment Over Time

Toryn Q. Klassen et al., Nov 2024

arXiv:2411.10654ProposalPartial

Birdwatch: Crowd Wisdom and Bridging Algorithms can Inform Understanding and Reduce the Spread of Misinformation

Stefan Wojcik et al., Oct 2022

arXiv:2210.15723PrecedentBuilt

A matrix-factorization algorithm picks which crowd-written annotations to show on a social media post by favoring those rated helpful by user groups that otherwise rate things differently, tested in a randomized survey experiment and in deployment on Twitter.

Supernotes: Driving Consensus in Crowd-Sourced Fact-Checking

Soham De et al., Nov 2024

arXiv:2411.06116PrecedentBuilt

A language model writes new fact-check notes by combining several existing community notes, and a scoring model trained on millions of past helpfulness ratings selects the candidate most likely to be rated helpful by a diverse set of users.

Everyone Conforms, No One Believes: Pluralistic Ignorance in LLM Agent Populations

Yashwanth YS, Aug 2026

arXiv:2608.02758BenchmarkBuilt

A benchmark of 100 scenarios across 10 domains and 5 authority levels measures how often language-model agents publicly go along with a norm they privately reject, and how often a single dissenting agent breaks the false consensus.

Aligned Alone, Misaligned Together: Forecasting Adversarial Capture in LLM Agent Populations

Isotta Magistrali & Chen Shani, Aug 2026

arXiv:2608.22444AnalysisBuilt

Populations of language-model monitors triage security alerts while a committed minority always pushes one way, and a response function calibrated on the population's behavior before any attack forecasts how far that minority will later move it.

SCENE: Recognizing Social Norms and Sanctioning in Group Chats

Mateusz Jacniacki & Maksymilian Bilski, May 2026

arXiv:2605.07823BenchmarkBuilt

A benchmark drops a model into a multi-party chat where scripted personas follow a hidden norm, create chances to break it and sanction the breach, then scores whether the model responds to the sanction and picks the norm up from its peers.

Customary pluralism

Graph Feedback Controls Consensus and Clique Formation in Open-Weight Language-Model Populations

Samer Saab & Chaouki Abdallah, Jul 2026

arXiv:2607.12077AnalysisBuilt

Runs a naming game across open-weight agents from 1.1B to 32B and shows that similarity-based routing can isolate an emerging convention and sustain fragmentation even when every agent interacts in every round, with matched controls ruling out uneven participation and model-family effects.

A theory of appropriateness with applications to generative artificial intelligence

Joel Z. Leibo et al., Dec 2024

arXiv:2412.19010ProposalGap

Sets out a theory of appropriateness — how the multi-scale mosaic of context-specific standards works in human society, how it might be implemented in the brain, and what follows for deploying generative AI responsibly.

A Theory of Appropriateness That Accounts for Norms of Rationality

Joel Z. Leibo et al., Mar 2026

arXiv:2603.140502 citationsProposalGap

Recasts appropriateness as pattern completion — each actor answering 'what does a person such as I do in a situation such as this?' — and shows this accounts for norms being context-dependent, arbitrary, automatic, dynamic and backed by sanction, against rational-choice accounts.

ValueScope: Unveiling Implicit Norms and Values via Return Potential Model of Social Interactions

Chan Young Park et al., Jul 2024

arXiv:2407.02472PrecedentBuilt

A framework uses language models to quantify the implicit norms and values of individual online communities from how their members write, applied to 13 Reddit communities grouped under gender, politics, science and finance.

CultureBank: An Online Community-Driven Knowledge Base Towards Culturally Aware Language Technologies

Weiyan Shi et al., Apr 2024

arXiv:2404.15238ResourceBuilt

A pipeline turns users' self-narratives from online platforms into a knowledge base of cultural descriptors, 12K sourced from TikTok and 11K from Reddit, which is then used both to evaluate models and to fine-tune one.

Values in the Wild: Discovering and Analyzing Values in Real-World Language Model Interactions

Saffron Huang et al., Apr 2025

arXiv:2504.15236AnalysisBuilt

A privacy-preserving method extracts the values a model states or demonstrates across hundreds of thousands of real-world interactions and organizes them into a taxonomy of 3,307 values.

CulturePark: Boosting Cross-cultural Understanding in Large Language Models

Cheng Li et al., May 2024

arXiv:2405.15145MethodBuilt

Language-model agents play people from different cultures and talk to one another, and the 41,000 samples the dialogues produce are used to fine-tune eight culture-specific models.

NormAd: A Framework for Measuring the Cultural Adaptability of Large Language Models

Abhinav Rao et al., Apr 2024

arXiv:2404.12464BenchmarkBuilt

An evaluation framework asks a model whether a described situation is socially acceptable while varying how explicitly the relevant cultural norm is supplied, instantiated as a set of 2.6k situational descriptions covering social-etiquette norms from 75 countries.

STEER-BENCH: A Benchmark for Evaluating the Steerability of Large Language Models

Kai Chen et al., May 2025

arXiv:2505.20645BenchmarkBuilt

A benchmark built from 30 contrasting subreddit pairs across 19 domains tests whether a model can answer in line with one named community's norms, using over 10,000 instruction-response pairs and 5,500 validated multiple-choice questions with silver labels.

CCD-Bench: Probing Cultural Conflict in Large Language Model Decision-Making

Hasibur Rahman & Hanan Salam, Oct 2025

arXiv:2510.03553BenchmarkBuilt

A benchmark of 2,182 open-ended dilemmas across seven domains asks a model to choose between ten anonymized responses, each corresponding to one of the ten GLOBE cultural clusters, so its default cultural preference can be read off the choices.

Diverse Conventions for Human-AI Collaboration

Bidipta Sarkar et al., Oct 2023

arXiv:2310.15414MethodBuilt

Trains a collection of agents that each learn a different way of coordinating in a cooperative game, by rewarding an agent for playing well with copies of itself and badly with the ways of coordinating already found.

The CARE Principles for Indigenous Data Governance

Stephanie Russo Carroll et al., Nov 2020

doi:10.5334/dsj-2020-043928 citationsPrecedentBuilt

Sets out four principles for Indigenous data governance, collective benefit, authority to control, responsibility and ethics, which govern live data repositories.

Character alignment

Claude's Character

Anthropic, Jun 2024

anthropic.comProposalBuilt

The Capacity for Moral Self-Correction in Large Language Models

Deep Ganguli et al. (Anthropic), Feb 2023

arXiv:2302.07459214 citationsAnalysisBuilt

Measures moral self-correction as a capacity that emerges with model scale and RLHF training.

Open Character Training: Shaping the Persona of AI Assistants through Constitutional AI

Sharan Maiya et al., Nov 2025

arXiv:2511.01689MethodBuilt

Releases the first open implementation of character training — Constitutional AI plus a synthetic-introspective-data pipeline — and fine-tunes three open-weights models to eleven example personas, from humorous to deeply caring to outright malevolent.

Whose Opinions Do Language Models Reflect?

Shibani Santurkar et al., Mar 2023

arXiv:2303.17548BenchmarkBuilt

Builds a dataset from public opinion polls and scores how closely a language model's answers to subjective questions match those of 60 US demographic groups.

Towards Measuring the Representation of Subjective Global Opinions in Language Models

Esin Durmus et al., Jun 2023

arXiv:2306.16388BenchmarkBuilt

Builds a dataset of questions from cross-national surveys and measures how similar a model's answers are to the answers people in each country actually gave.

Are Large Language Models Consistent over Value-laden Questions?

Jared Moore et al., Jul 2024

arXiv:2407.02996BenchmarkBuilt

Measures how far a model gives the same answer to value-laden questions across paraphrases of one question, related questions on one topic, multiple-choice and open-ended versions, and translations.

Evaluating the Moral Beliefs Encoded in LLMs

Nino Scherrer et al., Jul 2023

arXiv:2307.14324BenchmarkBuilt

Runs a survey of moral scenarios on 28 language models and reports, for each, the probability of the model choosing an action, the uncertainty attached to that choice and how consistent the choice is.

Phronetic adjudication

A Geometric Perspective on Stabilizing Value Conflict Resolution

Saket Reddy & Andy Liu, Jul 2026

arXiv:2607.17946MethodBuilt

Trains chain-of-thought reasoning aimed specifically at value conflicts and shows both that it smooths the loss landscape in its sharpest direction and that the resulting reasoning transfers to other kinds of moral judgment.

Are Language Models Consequentialist or Deontological Moral Reasoners?

Keenan Samway et al., May 2025

arXiv:2505.21479AnalysisBuilt

Reads the moral reasoning traces LLMs produce across more than 600 distinct trolley problems and classifies them against a taxonomy of rationales, finding that the chains of thought lean deontological — reasoning from moral obligations — while the post-hoc explanations lean consequentialist.

Normative Conflicts and Shallow AI Alignment

Raphaël Millière, Jun 2025

arXiv:2506.04679ProposalGap

Argues that fine-tuning on helpfulness, honesty and harmlessness installs shallow behavioral dispositions rather than the capacity to reason through conflicts between those norms, and that adversarial attacks succeed precisely by driving the norms against each other.

Automated Parliaments: A Solution to Decision Uncertainty and Misalignment in Language Models

Thomas Forster et al., Oct 2023

arXiv:2311.10098ProposalPartial

Sets out an architecture in which several AI delegates, each representing a perspective, generate responses aligned with their own theory, alter one another's responses to make them more self-aligned, and then collectively assess the best end response.

Reinforcement Learning Under Moral Uncertainty

Adrien Ecoffet & Joel Lehman, Jun 2020

arXiv:2006.04734MethodBuilt

Trains reinforcement learning agents whose credence is split across several plausible ethical theories, using two training methods that realize different points among competing desiderata, and observes how they behave in simple environments.

A Bargaining-Theoretic Approach to Moral Uncertainty

Hilary Greaves & Owen Cotton-Barratt, Dec 2023

doi:10.1163/17455243-202338106 citationsPrecedentGap

Treats rival moral theories as parties bargaining over a space of lotteries and proposes a Nash bargaining solution, with stated axioms and variants, as the rule for acting when one's credence is split between them.

Value Kaleidoscope: Engaging AI with Pluralistic Human Values, Rights, and Duties

Taylor Sorensen et al., Sep 2023

arXiv:2309.00779MethodBuilt

Given a situation, Kaleido lists the values, rights and duties that bear on it, says of each whether it supports or opposes the action and how relevant it is, and explains it in words.

Imagining and building wise machines: The centrality of AI metacognition

Samuel G. B. Johnson et al., Nov 2024

arXiv:2411.02478ProposalGap

It analyses human wisdom as two layers of strategy for problems that lie outside the scope of analytic techniques, object-level heuristics for managing problems and metacognitive strategies for managing those heuristics, and argues that AI systems particularly struggle with the second layer.

Steerable persona pluralism

Toward AI That Understands Self and Others: A World-Model Theory of Cognitive Diversity and Alignment

Toru Takahashi, May 2026

arXiv:2605.29930ProposalGap

Models each agent — human, AI or institution — as building approximate sufficient statistics under finite constraints, and defines alignment maps plus a transformation loss for passing content between two such world models without merging them.

The benefits, risks and bounds of personalizing the alignment of large language models to individuals

Hannah Rose Kirk et al., 2024

doi:10.1038/s42256-024-00820-y239 citationsProposalGap

Sets out what personalizing a model to an individual would buy, what it would cost, and where the bounds should sit — with the competing philosophical bases for drawing those bounds made explicit.

Modular Pluralism: Pluralistic Alignment via Multi-LLM Collaboration

Shangbin Feng et al., Jun 2024

arXiv:2406.15951MethodBuilt

A main LLM collaborating with a pool of community-specific LLMs across Overton, steerable and distributional modes.

Whose Alignment? Comparing LLM Process Alignment Across Diverse Organizational Decision Contexts

Niklas Weller & Emilio Barkett, May 2026

arXiv:2605.252560 citationsAnalysisPartial

The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models

Hannah Rose Kirk et al., Apr 2024

arXiv:2404.16019ResourceBuilt

Maps the demographics and stated preferences of 1,500 participants from 75 countries onto their actual feedback in 8,011 live conversations with 21 models, so a rating can be traced back to who gave it and what they said they valued.

Steerable Pluralism: Pluralistic Alignment via Few-Shot Comparative Regression

Jadie Adams et al., Aug 2025

arXiv:2508.08509MethodBuilt

Adapts to an individual user from a handful of their comparative judgments, using in-context learning over a set of fine-grained attributes rather than a single scalar reward.

ComPO: Community Preferences for Language Model Personalization

Sachin Kumar et al., Oct 2024

arXiv:2410.16027MethodBuilt

Conditions preference optimization on which community the preference came from, so one model produces different outputs for different communities instead of averaging their styles and norms away.

PERSONA: A Reproducible Testbed for Pluralistic Alignment

Louis Castricato et al., Jul 2024

arXiv:2407.17387BenchmarkBuilt

Procedurally generates 1,586 synthetic personas from US census data, elicits 317,200 feedback pairs from them across 3,868 prompts, and uses the result to test — against human judges — how well a model can role-play a stated user.

CommunityBench: Benchmarking Community-Level Alignment across Diverse Groups and Tasks

Jiayu Lin & Zhongyu Wei, Jan 2026

arXiv:2601.13669BenchmarkBuilt

Proposes community-level alignment as a middle ground between one-size-fits-all and per-individual customization, and builds the first large-scale benchmark for it — four tasks grounded in Common Identity and Common Bond theory.

Group Preference Optimization: Few-Shot Alignment of Large Language Models

Siyan Zhao et al., Oct 2023

arXiv:2310.11523MethodBuilt

Adds an independent transformer module to a base LLM that predicts a group's preferences over the model's generations from a few examples given in context, meta-learned across several groups, and tested on adapting to US demographic groups, to countries and to individual users.

Aligning to Thousands of Preferences via System Message Generalization

Seongyun Lee et al., May 2024

arXiv:2405.17977MethodBuilt

Trains a 7B model called Janus on 192k combinations of stated values spanning 65k user instructions so that what a user writes in the system message steers its behavior, tested on 921 prompts from five benchmarks under system messages it has not seen.

Cultivating Pluralism In Algorithmic Monoculture: The Community Alignment Dataset

Lily Hong Zhang et al., Jul 2025

arXiv:2507.09650ResourceBuilt

Runs a preference study with representative samples from five countries (N=15,000), finds that people vary far more in what they want than the responses of 21 state-of-the-art models do, and releases Community Alignment, a multilingual multi-turn dataset of 233,319 comparisons built by prompting for candidate answers that pull in opposite directions.

Evaluating the Prompt Steerability of Large Language Models

Erik Miehling et al., Nov 2024

arXiv:2411.12405BenchmarkBuilt

Defines steerability as how far a model's joint behavioral distribution can be shifted from its baseline by prompting, computes indices for that shift across persona dimensions and directions as steering effort rises, and releases the benchmark as running code.

Political Bias Audits of LLMs Capture Sycophancy to the Inferred Auditor

Petter Törnberg & Michelle Schimmel, Apr 2026

arXiv:2604.27633AnalysisBuilt

Administers the Political Compass Test, the Pew Political Typology and 1,540 partisan-benchmarked Pew American Trends Panel items to six frontier models while varying only the asker's stated identity (N = 30,990 responses), and finds that a conservative Republican cue moves all six models right of center and cuts the share of items closer to Democrats by 28 to 62 percentage points, while the mirrored progressive cue produces little change.

Deliberative aggregation

Democratic Inputs to AI (grant program)

OpenAI, May 2023

openai.comResourcePartial

Collective Constitutional AI: Aligning a Language Model with Public Input

Saffron Huang et al., Jun 2024

doi:10.1145/3630106.3658979204 citationsMethodBuilt

Uses Polis to crowd-source and vote on constitutional principles, fine-tunes on the result, and evaluates against a developer-written baseline.

Beyond Preferences in AI Alignment

Tan Zhi-Xuan et al., Aug 2024

arXiv:2408.1698471 citationsProposalGap

Names the "preferentist" commitments underlying mainstream alignment — that preferences represent values, that rationality is preference satisfaction, and that systems should be aligned to preferences — and sets out conceptual alternatives to each.

Position: Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Vincent Conitzer et al., 2024

arXiv:2404.10271108 citationsProposalPartial

Argues RLHF assumes a homogenized average preference and that social choice theory supplies the missing aggregation machinery.

Position: A Roadmap to Pluralistic Alignment

Taylor Sorensen et al., Feb 2024

arXiv:2402.05070221 citationsProposalPartial

Defines three operational forms of pluralism (Overton, steerable, distributional) plus three matching benchmark classes.

Representative Social Choice: From Learning Theory to AI Alignment

Tianyi Qiu, Oct 2024

arXiv:2410.239538 citationsProposalPartial

Using the Veil of Ignorance to align AI systems with principles of justice

Laura Weidinger et al., 2023

doi:10.1073/pnas.221370912048 citationsAnalysisPartial

N≈2,000 participants choose governing principles from behind a veil of ignorance and prioritize the worst-off.

Wide Reflective Equilibrium in LLM Alignment: Bridging Moral Epistemology and AI Safety

Matthew Brophy, May 2025

arXiv:2506.004150 citationsProposalPartial

Argues that the method of wide reflective equilibrium is the right description of what Constitutional AI already does, and proposes concrete ways to make its revision loop more legitimate.

Self-Improvement as Coherence Optimization: A Theoretical Account

Tianyi Qiu et al., Jan 2026

arXiv:2601.135661 citationMethodPartial

Shows that debate, bootstrapping and internal-coherence maximization are all instances of one thing — searching for the most compressible, jointly predictable mapping from context to behavior — and proves that this is equivalent to description-length regularization.

Reflective Verbal Reward Design for Pluralistic Alignment

Carter Blair et al., Jun 2025

arXiv:2506.178342 citationsMethodPartial

Walks each user through a reflective dialogue in which they critique agent behavior and build up their own preferences, then learns an individual reward model from that dialogue instead of one aggregate model.

Making Reflective Equilibrium Precise: A Formal Model

Claus Beisbart et al., 2021

doi:10.3998/ergo.1152PrecedentPartial

Builds an explicit formal model of reflective equilibrium in which commitments and principles are adjusted against each other under measurable account, systematicity and faithfulness criteria.

Chain of Alignment: Integrating Public Will with Expert Intelligence for Language Model Alignment

Andrew Konya et al., Nov 2024

arXiv:2411.105342 citationsMethodPartial

Democratic policy development using collective dialogues and AI

Andrew Konya et al., 2023

arXiv:2311.0224230 citationsMethodBuilt

Justifications for Democratizing AI Alignment and Their Prospects

André Steingrüber & Kevin Baum, Jul 2025

arXiv:2507.195483 citationsProposalGap

Weighs the instrumental case for democratic alignment (better outcomes) against the non-instrumental one (avoiding illegitimate authority), and asks which survives under normative uncertainty.

Democratizing value alignment: from authoritarian to democratic AI ethics

Linus Ta-Lun Huang et al., 2025

doi:10.1007/s43681-024-00624-113 citationsProposalGap

Argues that value alignment as practiced installs a small number of decision-makers as the authority on values, and sets out what a democratic alternative would have to look like.

Generative Social Choice

Sara Fish et al., Sep 2023

arXiv:2309.01291MethodPartial

Axioms for AI Alignment from Human Feedback

Luise Ge et al., May 2024

arXiv:2405.14758ProposalPartial

Distortion of AI Alignment: Does Preference Optimization Optimize for Preferences?

Paul Gölz et al., May 2025

arXiv:2505.2374920 citationsAnalysisPartial

A matter of principle? AI alignment as the fair treatment of claims

Iason Gabriel & Geoff Keeling, 2025

doi:10.1007/s11098-025-02300-416 citationsProposalGap

Benchmarking Overton Pluralism in LLMs

Elinor Poole-Dayan et al., Dec 2025

arXiv:2512.013519 citationsBenchmarkBuilt

Formalizes Overton pluralism as set coverage — how much of the range of defensible views a single answer contains — validates the metric against 1,208 US-representative human judgments, and finds current models cover only about a third to two-fifths of it.

AI can help humans find common ground in democratic deliberation

Michael Henry Tessler et al., Oct 2024

doi:10.1126/science.adq2852172 citationsMethodBuilt

Trains an AI mediator that takes a group's individual opinions and critiques, writes a candidate group statement, and refines it round after round; participants (N=5,734) preferred its statements to those of human mediators and moved toward a shared position, and the result replicated in a demographically representative UK citizens' assembly.

Finding Common Ground in a Sea of Alternatives

Jay Chooi et al., Mar 2026

arXiv:2603.16751MethodBuilt

Gives a formal account of finding common ground when the candidate statements are effectively infinite, based on the proportional veto core, and supplies a sampling algorithm that returns an alternative in the approximate core with high probability, with matching lower bounds.

Generating Fair Consensus Statements with Social Choice on Token-Level MDPs

Carter Blair & Kate Larson, Oct 2025

arXiv:2510.14106MethodBuilt

Models consensus-statement generation as a token-level Markov decision process with one objective per participant, deriving each participant's token rewards from their own personalized model, so fairness guarantees attach to the generation itself.

The Empty Chair: Using LLMs to Raise Missing Perspectives in Policy Deliberations

Suyash Fulay et al., Mar 2025

arXiv:2503.1381210 citationsApplicationBuilt

Transcribes a live deliberation and injects, in real time, the contributions that relevant but absent stakeholders would have made, using LLM personas; deployed in a 19-person student deliberation.

Generative Social Choice: The Next Generation

Niclas Boehmer et al., May 2025

arXiv:2505.22939MethodBuilt

Extends generative social choice to produce a slate of statements that proportionally represents the whole spectrum of opinion, with theoretical guarantees that hold under the queries an LLM can actually answer.

Jackpot! Alignment as a Maximal Lottery

Roberto-Rafael Maura-Rivero et al., Jan 2025

arXiv:2501.19266MethodBuilt

Replaces RLHF's reward maximization with maximal lotteries — a randomized social-choice rule — and shows that a family of alignment methods including Nash learning already sits inside that framework.

MaxMin-RLHF: Alignment with Diverse Human Preferences

Souradip Chakraborty et al., Feb 2024

arXiv:2402.08925MethodBuilt

Proves that a single reward model cannot represent diverse human preferences, then learns a mixture of reward models and optimizes the worst-off group's reward — an egalitarian objective in place of the average.

Nash Learning from Human Feedback

Rémi Munos et al., Dec 2023

arXiv:2312.00886MethodBuilt

Drops the reward model and instead learns a preference model, then seeks the policy whose responses are preferred to those of any competing policy — the Nash equilibrium of that preference model — with a mirror-descent algorithm (Nash-MD) that converges to the regularized equilibrium.

Policy Aggregation

Parand A. Alamdari et al., Nov 2024

arXiv:2411.03651MethodBuilt

Formalizes aligning to several people at once as aggregating their optimal policies rather than their reward functions, and shows social-choice rules can be applied by identifying ordinal preferences with volumes of state-action space.

AI Alignment and Social Choice: Fundamental Limitations and Policy Implications

Abhilash Mishra, Oct 2023

arXiv:2310.16048ProposalGap

Builds on impossibility results in social choice to show that no voting protocol can universally align AI systems through RLHF, and that aligning to everyone's values must violate some individual's private ethical preferences.

Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework

Kihyun Kim et al., Jun 2025

arXiv:2506.05619MethodBuilt

Aligns policies proportionally to the true distribution of evaluator preferences rather than to whichever opinion is held most widely, under an axiomatic framework that also makes the result harder to manipulate strategically.

AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence

Paul Anton Bachmann et al., Apr 2026

arXiv:2605.16291ProposalGap

Argues that collective control of AI should enter at many points across the development pipeline — data, objectives, alignment, deployment — rather than only at macro-level governance, and works through what social choice offers at each point.

Democratic Preference Alignment via Sortition-Weighted RLHF

Suvadip Sana et al., Feb 2026

arXiv:2602.05113MethodBuilt

Draws the rater pool by algorithmic sortition — the same lottery mechanism used to seat citizens' assemblies — and offers both a hard-panel scheme that trains only on the drawn panel and a soft-weighting scheme, so representativeness is built into the training data rather than corrected afterwards.

What are human values, and how do we align AI to them?

Oliver Klingefjord et al., Mar 2024

arXiv:2404.10636MethodBuilt

Has a language model interview participants about the values they actually brought to a hard, divisive question, then has them judge which value is wiser in that context — building a moral graph whose winners become the alignment target; trialled with a representative sample of 500 Americans on three divisive prompts.

Deliberative Technology for Alignment

Andrew Konya et al., Dec 2023

arXiv:2312.03893ProposalGap

Argues that the deliberative technology already used by governments, firms and NGOs is the natural substrate for aligning AI with collective will, and sets out what it would take to scale it to that job.

Fine-tuning language models to find agreement among humans with diverse preferences

Michiel A. Bakker et al., Nov 2022

arXiv:2211.15006MethodBuilt

Fine-tunes a 70-billion-parameter model to write statements that maximize a group's expected approval and ranks the candidates with a reward model trained to predict individual preferences, with the group's appeal defined according to different social welfare functions; its statements were preferred to those of prompted models more than 70% of the time and to the best human-written opinions more than 65%, and when a consensus was built silently from only a subset of the group the excluded members were more likely to dissent.

Can AI mediation improve democratic deliberation?

Michael Henry Tessler et al., Jan 2026

arXiv:2601.05904ProposalGap

A discussion paper that takes the LLM mediation system of Tessler, Bakker and colleagues and asks how it bears on Fishkin's trilemma between broad participation, meaningful deliberation and political equality, setting out where scalability, fair mediation and the surfacing of trustworthy information might help and where challenges remain.

Constitutional Governance in Metric Spaces

Ehud Shapiro & Nimrod Talmon, May 2026

arXiv:2605.13362PrecedentGap

Sets out a voting protocol in which a community's laws and its constitution are each a point in a metric space: members submit their ideal points, a polynomial-time rule scores proposals that already carry supermajority public support, and a proposal whose score is positive and maximal for two rounds running is adopted, otherwise the status quo is retained.

Using Collective Dialogues and AI to Find Common Ground Between Israeli and Palestinian Peacebuilders

Andrew Konya et al., Mar 2025

arXiv:2503.01769ApplicationBuilt

Runs an iterative deliberative process from April to July 2024 with around 138 Israeli and Palestinian civil society peacebuilders, combining large language models, bridging-based ranking and collective dialogues, and produces collective statements including demands to world leaders with at least 84% agreement from participants on each side.

Achieving parity with human moderators

Lodewijk Gelauff et al., Jun 2023

doi:10.4324/9781003215929-156 citationsPrecedentBuilt

Describes the Stanford Online Deliberation Platform, whose automated moderator manages the speaking queue, allocates speaking time, moves the agenda along and intervenes on incivility, and reports it reaching parity with human moderators across real Deliberative Polling sessions.

Strategyproof Reinforcement Learning from Human Feedback

Thomas Kleine Buening et al., Mar 2025

arXiv:2503.09561MethodPartial

Shows that existing RLHF methods, pluralistic ones included, can be pushed arbitrarily far from social welfare by a single labeler who misreports, proves that any strategyproof rule must in the worst case do k times worse than the optimal policy when there are k labelers, and gives a Pessimistic Median of MLEs algorithm that is approximately strategyproof and converges to the optimum as labelers and samples grow.

Soft Condorcet Optimization for Ranking of General Agents

Marc Lanctot et al., Oct 2024

arXiv:2411.00119MethodBuilt

Ranks agents by treating benchmark and tournament results as votes and searching for the ranking that mispredicts the fewest pairwise comparisons, landing on average 0 to 0.043 in normalized Kendall-tau from the optimal ranking across 865 PrefLib preference profiles and giving the best approximation to the optimal ranking on held-out test sets from 31,049 games of seven-player Diplomacy played by 52,958 people.

Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models

Pat Verga et al., Apr 2024

arXiv:2404.18796BenchmarkBuilt

Scores model outputs with a panel of several smaller models drawn from disjoint model families instead of one large judge, and reports better performance than a single large judge, less intra-model bias and over seven times lower cost across three judge settings and six datasets.

Adaptive Pluralistic Alignment: A pipeline for dynamic artificial democracy

Rachel Freedman, May 2026

arXiv:2605.01642MethodPartial

Learns a compact reward model per annotator by low-rank decomposition over a shared reward basis, has those models vote as a jury over candidate outputs under a social-choice rule, and adapts the jury over time by fitting new annotator weights on the fixed bases as values shift; a proof-of-concept on the PRISM dataset with simulated historical annotators finds that jury composition and the choice of voting rule can substantially affect outcomes where jury preferences are heterogeneous.

Human-centred mechanism design with Democratic AI

Raphael Koster et al., Jul 2022

doi:10.1038/s41562-022-01383-x72 citationsMethodBuilt

A deep reinforcement learning pipeline designs the redistribution mechanism for a multiplayer public-goods game by optimizing it to be preferred by majority vote, first against simulated players and then in elections among real people.

Eliciting Human Preferences with Language Models

Belinda Z. Li et al., Oct 2023

arXiv:2310.11589MethodBuilt

A language model interviews the user by generating open-ended questions or synthesizing informative edge cases, then infers from the answers what behavior the user intended.

Active Preference-Based Learning of Reward Functions

Dorsa Sadigh et al., Jul 2017

doi:10.15607/rss.2017.xiii.053168 citationsMethodBuilt

An algorithm generates the pair of trajectories to put to a person next, choosing the comparison expected to be most informative, and fits a reward function to the answers.

AQuA -- Combining Experts' and Non-Experts' Views To Assess Deliberation Quality in Online Discussions Using LLMs

Maike Behrendt et al., Apr 2024

arXiv:2404.02761PrecedentBuilt

AQuA gives each post in a political online discussion a single deliberative quality score by adding the outputs of adapter models for 20 deliberative indices, each index weighted by the correlation between experts' annotations of it and non-experts' perceived deliberativeness.

Estimating Contribution Quality in Online Deliberations Using a Large Language Model

Lodewijk Gelauff et al., Aug 2024

arXiv:2408.11936PrecedentBuilt

A large language model rates each contribution in a video deliberation from 1 to 5 on justification, novelty, expansion of the conversation and potential for further expansion, doing the work human annotators had been doing.

Adversarial proceduralism

AI Safety via Debate

Geoffrey Irving et al., May 2018

arXiv:1805.00899432 citationsMethodBuilt

Two agents argue before a human judge; the adversarial structure lets a weaker judge evaluate stronger agents.

Scalable Agent Alignment via Reward Modeling: A Research Direction

Jan Leike et al., Nov 2018

arXiv:1811.07871617 citationsProposalBuilt

Recursive reward modeling: train a separate network to stand in for human judgment, then optimize against it.

Negotiative Alignment: Embracing Disagreement to Achieve Fairer Outcomes — Insights from Urban Studies

Rashid Mushkani et al., Mar 2025

arXiv:2503.1261310 citationsPrecedentPartial

Community study (n=35) plus a budget-aware bargaining procedure with role-played stakeholders and a neutral mediator.

Arbiters of Ambivalence: Challenges of Using LLMs in No-Consensus Tasks

Bhaktipriya Radharapu et al., May 2025

arXiv:2505.238208 citationsBenchmarkPartial

On scalable oversight with weak LLMs judging strong LLMs

Zachary Kenton et al., Jul 2024

arXiv:2407.04622101 citationsAnalysisBuilt

Runs debate, consultancy and plain question-answering against each other with weaker models standing in for human judges, across information, mathematics, coding, logic and multimodal asymmetries: debate beats consultancy on every task, and beats direct question-answering only where the judge lacks information the debaters have.

Training Language Models to Win Debates with Self-Play Improves Judge Accuracy

Samuel Arnesen et al., Sep 2024

arXiv:2409.16636MethodBuilt

Trains debaters by self-play to win, and shows that judges get more accurate as the debaters get better — while models trained to persuade without an opponent present produce no such gain.

AI Debate Aids Assessment of Controversial Claims

Salman Rahman et al., Jun 2025

arXiv:2506.02175AnalysisBuilt

Puts two AI systems to debate opposing sides of contested COVID-19 and climate claims in front of human judges holding mainstream or skeptical prior beliefs, and finds debate beats a single AI consultant by 4–10% on accuracy — up to +15.2% for mainstream judges, and +4.7% even for skeptical judges who initially had the claim wrong.

Prover-Verifier Games improve legibility of LLM outputs

Jan Hendrik Kirchner et al., Jul 2024

arXiv:2407.13692MethodBuilt

Iteratively trains a small verifier to predict whether a solution is correct, a helpful prover to produce correct solutions the verifier accepts, and a sneaky prover to produce incorrect ones that fool it — and shows the resulting legibility transfers to time-constrained humans, whose accuracy rises on the helpful prover's solutions and falls on the sneaky prover's.

Preserving Disagreement: Architectural Heterogeneity and Coherence Validation in Multi-Agent Policy Simulation

Ariel Sela, Apr 2026

arXiv:2604.26561AnalysisBuilt

Runs a three-phase deliberating council over 120 deliberations and shows that giving each value perspective its own 7–9B model cuts first-choice concentration sharply (70.9% to 46.1% on child welfare, 46.0% to 22.9% on housing), where accuracy-oriented multi-agent debate gets no such benefit from model diversity.

Self-critiquing models for assisting human evaluators

William Saunders et al., Jun 2022

arXiv:2206.05802MethodBuilt

Language models are fine-tuned by behavioral cloning to write natural-language critical comments on summaries, including their own, which human evaluators read and which larger models can fold back into a revised summary.

Scalable Oversight for Superhuman AI via Recursive Self-Critiquing

Xueru Wen et al., Feb 2025

arXiv:2502.04675AnalysisPartial

Higher-order critiques, a critique of a critique and a critique of that, are compared for difficulty against the level beneath them in human-human, human-AI and AI-AI experiments.

Supervising strong learners by amplifying weak experts

Paul Christiano et al., Oct 2018

arXiv:1810.08575MethodBuilt

Iterated Amplification trains a model on problems too complicated for a human to evaluate directly by progressively building a training signal out of solutions to easier subproblems, and reports results in algorithmic environments.

Debating with More Persuasive LLMs Leads to More Truthful Answers

Akbir Khan et al., Feb 2024

arXiv:2402.06782AnalysisBuilt

Two stronger models that hold the information needed to answer each argue for a different answer while a weaker model or a human who lacks that information picks between them, with debate reaching 76% accuracy for the model judges and 88% for the human judges against naive baselines of 48% and 60%.

Debate Helps Supervise Unreliable Experts

Julian Michael et al., Nov 2023

arXiv:2311.08702AnalysisBuilt

Collects human-written debates on hard reading comprehension questions where the judge has not read the source passage and sees only the arguments and the short quotes the debaters selectively reveal, and finds judges reach 84% accuracy under debate against 74% under a single consultant, with debates running 68% of the length.

Scalable AI Safety via Doubly-Efficient Debate

Jonah Brown-Cohen et al., Nov 2023

arXiv:2311.14125ProposalGap

Sets out a new set of debate protocols in which the honest side can always succeed using a simulation of polynomially many steps and can verify the alignment of stochastic AI systems, even where the dishonest side is allowed exponentially many simulation steps.

Avoiding Obfuscation with Prover-Estimator Debate

Jonah Brown-Cohen et al., Jun 2025

arXiv:2506.13609ProposalGap

A recursive debate protocol pairing a prover with an estimator, put up against the case where a dishonest debater decomposes a problem into subproblems that force an honest opponent to solve something computationally intractable, and shown to let the honest debater win under certain stability assumptions with a strategy whose computational cost is comparable to its opponent's.

Neural Interactive Proofs

Lewis Hammond & Sam Adam-Day, Dec 2024

arXiv:2412.08897MethodBuilt

Sets out a family of prover-verifier games in which a trusted but computationally bounded verifier learns to interact with one or more powerful untrusted provers, compares new and existing protocols theoretically, and tests them on a toy graph isomorphism problem and on a code validation task using large language models.

Learning to Give Checkable Answers with Prover-Verifier Games

Cem Anil et al., Aug 2021

arXiv:2108.12099MethodBuilt

Sets up a game between a trusted verifier network trying to choose the correct answer and a more powerful but untrusted prover network trying to persuade it of a particular answer regardless of that answer's correctness, narrows the variants to a subset that provably has the desired equilibria, and shows on two algorithmic tasks that the verifier learns a robust decision rule which still works when the verifier is frozen and the prover's messages are optimized directly to convince it.

Collaborative Disagreement Resolution for Scalable Oversight

Yuyang Jiang et al., Jun 2026

arXiv:2607.01251MethodBuilt

An automated pipeline directs models to identify the points on which they disagree, examine the evidence for the conflicting claims and either converge on consensus or isolate the specific crux of the disagreement, in place of arguing fixed opposing positions in front of a judge.

Statistical Impossibility and Possibility of Aligning LLMs with Human Preferences: From Condorcet Paradox to Nash Equilibrium

Kaizhao Liu et al., Mar 2025

arXiv:2503.10990AnalysisGap

Shows that a reward model can represent human preferences over a model's answers exactly when those preferences contain no majority cycle, that such cycles arise with probability converging to one exponentially fast under the Luce model of choice, and that an alignment method using no reward model, taken in the limit, keeps several answers in play whenever no answer is preferred over all others by a majority.

Opportunities and Risks of LLMs for Scalable Deliberation with Polis

Christopher T. Small et al., Jun 2023

arXiv:2306.11932ApplicationPartial

Reports pilot experiments in which Anthropic's Claude is used to help facilitate, moderate and summarize conversations on Polis, a platform that uses machine intelligence to scale up deliberative processes, and discusses the risks that come with putting a model in that role.

AI and the Future of Digital Public Squares

Beth Goldberg et al., Dec 2024

arXiv:2412.09988ProposalGap

Sets out four ways language models could be used in online public discussion, namely collective dialogue systems, bridging systems, community moderation and proof-of-humanity systems, together with the risks they pose and an agenda for further research and investment.

Polycentric proceduralism

Decentralising LLM Alignment: A Case for Context, Pluralism, and Participation

Oriane Peter & Kate Devlin, Sep 2025

arXiv:2509.088589 citationsProposalPartial

Argues from the power/knowledge nexus that current alignment centralizes control over knowledge production.

Density-Guided Response Optimization: Community-Grounded Alignment via Implicit Acceptance Signals

Patrick Gerard & Svitlana Volkova, Mar 2026

arXiv:2603.032420 citationsMethodPartial

Learns each online community’s norms from what it already accepts and engages with, so no explicit preference labels or written principles are needed.

PluralLLM: Pluralistic Alignment in LLMs via Federated Learning

Mahmoud Srewa et al., Mar 2025

arXiv:2503.0992517 citationsMethodBuilt

Lets several user groups train a shared transformer preference predictor by federated averaging without any group's feedback leaving it, converging 46% faster than centralized training with a 4% higher alignment score and near-identical group fairness.

Towards Federated RLHF with Aggregated Client Preference for LLMs

Feijie Wu et al., Jul 2024

arXiv:2407.03038MethodBuilt

Encodes each client's preferences as binary selectors and aggregates the selectors rather than the data, grouping clients with similar preferences to handle heterogeneity and using several selectors at once to resist reward hacking.

FedRLHF: A Convergence-Guaranteed Federated Framework for Privacy-Preserving and Personalized RLHF

Flint Xiaofeng Fan et al., Dec 2024

arXiv:2412.15538MethodBuilt

Each client folds its own human feedback into a local reward function and updates its own policy through a personalized RLHF loop, with no raw data or human feedback leaving the client, reaching performance on a par with centralized RLHF on the MovieLens and IMDb datasets while improving personalization across client environments, and carrying convergence guarantees and sample complexity bounds that scale efficiently with the number of clients.

RLHF from Heterogeneous Feedback via Personalization and Preference Aggregation

Chanwoo Park et al., Apr 2024

arXiv:2405.00254MethodBuilt

Sets out two ways of handling human feedback that is not homogeneous, one learning several reward models by representation learning or by clustering, the other keeping the single-model pipeline and aggregating either the individual reward models under utilitarian and Leximin rules or the feedback itself as probabilistic opinions, with sample complexity guarantees for the personalization approaches and for reward aggregation, and a mechanism-design step that ensures truthful preference reporting by strategic labelers.

Collaborative Content Moderation in the Fediverse

Haris Bin Zia et al., Jan 2025

arXiv:2501.05871MethodBuilt

Lets Fediverse servers exchange the parameters of their partially trained local moderation models with similar servers to form a federated model shared among the collaborating servers, reaching average per-server macro-F1 of 0.71 on harmful content detection, 0.73 on bot content detection and 0.58 on content warning assignment.

PolicyKit: Building Governance in Online Communities

Amy X. Zhang et al., Oct 2020

doi:10.1145/3379337.341585863 citationsPrecedentBuilt

Open-source software, deployed on Reddit and Slack communities, in which a community writes its own governance procedures as Python policies attached to platform actions and a shared runtime executes them.

Modular Politics: Toward a Governance Layer for Online Communities

Nathan Schneider et al., May 2020

arXiv:2005.13701PrecedentGap

Sets out a design in which online-community governance is built bottom-up from software components that are modular, composable, portable from one context to another and interoperable across platforms, so that features absent from platform software such as juries, political parties, term limits and formal debates could be implemented, and calls for an open standard for networked governance.

Bottom-up data Trusts: disturbing the ‘one size fits all’ approach to data governance

Sylvie Delacroix & Neil D Lawrence, Oct 2019

doi:10.1093/idpl/ipz01487 citationsPrecedentGap

Sets out a legal structure in which many data trusts each hold data rights under their own trust deed and fiduciary terms, with the member's operative right being to leave one trust for another.

Data Cooperatives: Towards a Foundation for Decentralized Personal Data Management

Thomas Hardjono & Alex Pentland, May 2019

arXiv:1905.08819PrecedentGap

Describes cooperatives that hold their citizen members' personal data under a fiduciary obligation, manage, curate and protect access to it, run internal analytics to obtain insights about members' well-being, and use those insights to negotiate better services and discounts for them.

Data Governance in the Age of Large-Scale Data-Driven Language Technology

Yacine Jernite et al., May 2022

arXiv:2206.03216ProposalGap

Proposes an approach to global language data governance that organizes data management among stakeholders, values and rights as a multi-party international governance structure, and names the technical and organizational tools it would need to do its work.

Surveys, background & counter-positions

Political studies of automated governing: A bird's eye (re)view

Andreas Öjehag-Pettersson et al., Dec 2023

doi:10.1111/rego.125693 citationsAnalysis

AI as Governance

Henry Farrell, Jun 2025

doi:10.1146/annurev-polisci-040723-0132458 citationsProposal

Terra Incognita: The Governance of Artificial Intelligence in Global Perspective

Allison Stanger et al., Jul 2024

doi:10.1146/annurev-polisci-041322-04224711 citationsAnalysis

Large Language Model Alignment: A Survey

Tianhao Shen et al., Sep 2023

arXiv:2309.15025328 citationsAnalysis

AI Alignment: A Comprehensive Survey

Jiaming Ji et al., Oct 2023

arXiv:2310.19852382 citationsAnalysis

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

Haoran Lu et al., Jul 2025

arXiv:2507.1967215 citationsAnalysis

The Digital Leviathan: Hobbes Sovereignty in Digital Times

Hongtian Rang, Apr 2026

doi:10.2139/ssrn.6991063Proposal

Position: Align AI to Our Aspirations, Not Our Flaws

Nikita Kazeev et al., Jun 2026

arXiv:2606.137550 citationsProposal

Argues for a non-negotiable objective floor with pluralism confined above it, pre-engaging six objections.

AI Alignment From Social Choice Perspectives

Daniel Halpern et al., Jun 2026

arXiv:2606.21550Analysis

Surveys the body of work that reads alignment from human feedback as a preference-aggregation problem, and sets out the failure modes the social-choice lens exposes in the feedback layer.

AI Pluralism and the Worlds It Misses

Rashid Mushkani, Jun 2026

arXiv:2606.16167Proposal

Argues that framing pluralism as representing diverse values misses a prior imposition: AI systems fix what counts as an entity, a harm, a benefit and valid evidence at all, flattening situated meanings into technical categories treated as neutral.