Taxonomy / Deliberative aggregation / Writing a voting rule into the training objective

Writing a voting rule into the training objective

The standard pipeline pools everyone's comparisons into one reward model, which quietly averages disagreeing people into a single stand-in judge. This family makes an explicit collective decision rule the thing the system optimizes instead: maximize the worst-off group's reward, find the policy that beats every rival head to head, draw from a maximal lottery, aggregate whole policies rather than reward numbers, match the population's actual distribution of views, or — in the limiting case, where the object designed is an economic redistribution rule rather than a language policy — optimize directly to win an election among the people the rule governs. The same machinery is turned back on itself to say what any such rule can deliver: the axioms it satisfies, how far it can fall short of the ideal, whether one labeler who misreports can wreck it, and impossibility results showing that no protocol satisfies everyone.

The method, against Deliberative aggregation

Scroll the diagram sideways to see all of it.

Concept Analysis: Theoretical Foundations

Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.

Deliberation, Not Tallying

AbsentNot claimed · out of scope

Definition · Jürgen Habermas, Between Facts and Norms (1992)

This concept locates legitimacy in the exchange of reasons among equals: participants justify their views to each other, free of coercion, and remain able to change their minds. Counting votes or averaging preferences does not produce it, because the reasoning is what does the work.

Analysis

Every member of this family swaps one tallying rule for another — a maximal lottery (NEW-S36), the Nash equilibrium of a preference model (NEW-S38), maximin over a mixture of reward models (NEW-S37), proportional representation of the evaluator distribution (NEW-S41), a pessimistic median of MLEs (NEW-S154), or a literal majority ballot in Democratic AI (NEW-S158). All of them run on judgments already harvested from people who never meet each other, and no stage exists at which anyone defends a comparison, hears an objection, or revises a rating. Democratic AI is the sharpest case: real humans do vote, and the vote is the whole of the legitimacy claim, which is precisely the substitution Habermas denies. The analytical papers (NEW-S40, NEW-I3, NEW-S154) study properties of counting rules, so they inherit the same restriction.

General Will vs. Sum of Preferences

AbsentNot claimed · out of scope

Definition · Jean-Jacques Rousseau, The Social Contract (1762)

This distinction separates what is good for a public in common from the sum of what its members privately want. Aggregating private wants, however fairly, does not produce the former, since the two can and often do diverge.

Analysis

The input to every rule here is a private want — a labeler's comparison between two outputs, or in NEW-S158 a player's judgment of a redistribution rule that sets his own payoff — and the papers differ only in how those wants are weighted: worst-off, Condorcet, proportional, median. Rousseau's distinction turns on the input rather than the weighting, since the general will requires citizens to judge what is good in common, and nothing in the pipeline asks for that judgment or could tell it apart from self-interest. NEW-S41 is the clearest illustration: proportionality is a promise to mirror the population's private distribution faithfully. NEW-S40 proves a formal cousin of the gap — aligning to everyone must override some individual's private ethics — but treats it as an impossibility for tallying, not as a reason to seek a different object.

Veil of Ignorance

PartialClaimed · partial

Definition · John Rawls, A Theory of Justice (1971)

This device requires that rules be chosen without knowledge of which position the chooser will occupy under them, e.g., rich or poor, majority or minority. Not knowing generally pushes the chooser to protect the worst-off position, since it may turn out to be their own.

Analysis

MaxMin-RLHF writes the veil's characteristic output straight into the loss: it clusters raters into groups and maximizes the worst-off group's reward, so a majority's mild gains cannot buy a minority's loss. What is entirely missing is the device itself — the maximin objective is imposed by the designer, not reached by choosers deprived of knowledge of their position, and every rater in every paper reports from their own standpoint with their group membership inferred rather than hidden. Democratic AI runs the inversion: players vote on redistribution rules knowing exactly what each rule pays them, and an egalitarian mechanism is one of the baselines their self-interested votes defeat. Claimed is IMPLIED because MaxMin-RLHF sells its objective in the egalitarian-justice vocabulary the veil underwrites, and NEW-S158 names a Rawlsian-style baseline, while no source asserts position-blind choice.

Reasonable Rejection

PartialNot claimed

Definition · T. M. Scanlon, What We Owe to Each Other (1998)

This test holds a principle justified only if no individual could reasonably reject it, and it is applied person by person rather than in aggregate. One sufficiently strong objection therefore outweighs many mild preferences, which is the case averaging handles wrongly.

Analysis

MaxMin-RLHF has the structural shape Scanlon's test demands: the group with the lowest reward sets the objective, so no accumulation of mild approvals can outweigh it. Two things are missing. The unit is a latent cluster discovered by fitting a mixture, not a person, and the strength of an objection enters only as a scalar reward value rather than as a reason anyone could state and have answered; nobody is given standing to reject anything. And the rest of the family runs the other way — maximal lotteries, Nash equilibria, proportional alignment and Democratic AI's majority-vote training objective all let many weak preferences beat one severe complaint by construction, which NEW-S154 sharpens by showing how far a single participant's report can be made to matter or not matter at all.

Concept Analysis: Newly Introduced

Impossibility results as design constraints

Added

Social choice theory brings formal limits with it. Arrow's theorem shows that no voting rule can turn individual rankings into a group ranking while meeting a small set of basic fairness conditions at once. These works import that result into alignment, which establishes what aggregation cannot be expected to deliver.

The institution itself as the optimization target

Added

Democratic theory has constitutional design and it has ratification, but it assumes a rule is drafted by someone and then put to the people. Here the rule is searched for by reinforcement learning with the vote as the objective function, and the electorate is simulated during training so that candidate mechanisms can be tested against modelled players before real ones vote. Neither Rousseau nor Rawls nor Habermas describes a legislator that proposes and discards institutions at that rate, or a public modelled closely enough to be optimized against.

Papers

Jackpot! Alignment as a Maximal Lottery

Roberto-Rafael Maura-Rivero et al., Jan 2025

arXiv:2501.19266MethodBuilt

Replaces RLHF's reward maximization with maximal lotteries — a randomized social-choice rule — and shows that a family of alignment methods including Nash learning already sits inside that framework.

MaxMin-RLHF: Alignment with Diverse Human Preferences

Souradip Chakraborty et al., Feb 2024

arXiv:2402.08925MethodBuilt

Proves that a single reward model cannot represent diverse human preferences, then learns a mixture of reward models and optimizes the worst-off group's reward — an egalitarian objective in place of the average.

Nash Learning from Human Feedback

Rémi Munos et al., Dec 2023

arXiv:2312.00886MethodBuilt

Drops the reward model and instead learns a preference model, then seeks the policy whose responses are preferred to those of any competing policy — the Nash equilibrium of that preference model — with a mirror-descent algorithm (Nash-MD) that converges to the regularized equilibrium.

Policy Aggregation

Parand A. Alamdari et al., Nov 2024

arXiv:2411.03651MethodBuilt

Formalizes aligning to several people at once as aggregating their optimal policies rather than their reward functions, and shows social-choice rules can be applied by identifying ordinal preferences with volumes of state-action space.

Beyond RLHF and NLHF: Population-Proportional Alignment under an Axiomatic Framework

Kihyun Kim et al., Jun 2025

arXiv:2506.05619MethodBuilt

Aligns policies proportionally to the true distribution of evaluator preferences rather than to whichever opinion is held most widely, under an axiomatic framework that also makes the result harder to manipulate strategically.

Strategyproof Reinforcement Learning from Human Feedback

Thomas Kleine Buening et al., Mar 2025

arXiv:2503.09561MethodPartial

Shows that existing RLHF methods, pluralistic ones included, can be pushed arbitrarily far from social welfare by a single labeler who misreports, proves that any strategyproof rule must in the worst case do k times worse than the optimal policy when there are k labelers, and gives a Pessimistic Median of MLEs algorithm that is approximately strategyproof and converges to the optimum as labelers and samples grow.

Human-centred mechanism design with Democratic AI

Raphael Koster et al., Jul 2022

doi:10.1038/s41562-022-01383-x72 citationsMethodBuilt

A deep reinforcement learning pipeline designs the redistribution mechanism for a multiplayer public-goods game by optimizing it to be preferred by majority vote, first against simulated players and then in elections among real people.

Findings about it

What existing systems were observed to do. These characterise the problem the approach addresses without introducing a mechanism.

Argued for, not built

These make the case for this approach, or sketch a design for it, but leave nothing built and tested behind.

Position: Social Choice Should Guide AI Alignment in Dealing with Diverse Human Feedback

Vincent Conitzer et al., 2024

arXiv:2404.10271108 citationsProposalPartial

Argues RLHF assumes a homogenized average preference and that social choice theory supplies the missing aggregation machinery.

Position: A Roadmap to Pluralistic Alignment

Taylor Sorensen et al., Feb 2024

arXiv:2402.05070221 citationsProposalPartial

Defines three operational forms of pluralism (Overton, steerable, distributional) plus three matching benchmark classes.

Representative Social Choice: From Learning Theory to AI Alignment

Tianyi Qiu, Oct 2024

arXiv:2410.239538 citationsProposalPartial

Axioms for AI Alignment from Human Feedback

Luise Ge et al., May 2024

arXiv:2405.14758ProposalPartial

AI Alignment and Social Choice: Fundamental Limitations and Policy Implications

Abhilash Mishra, Oct 2023

arXiv:2310.16048ProposalGap

Builds on impossibility results in social choice to show that no voting protocol can universally align AI systems through RLHF, and that aligning to everyone's values must violate some individual's private ethical preferences.

AI of the People, by the People, for the People: A Social Choice Approach to Collective Control of Artificial Intelligence

Paul Anton Bachmann et al., Apr 2026

arXiv:2605.16291ProposalGap

Argues that collective control of AI should enter at many points across the development pipeline — data, objectives, alignment, deployment — rather than only at macro-level governance, and works through what social choice offers at each point.