Taxonomy / Deliberative aggregation

Deliberative aggregation

No norms are fixed in advance; a procedure runs at the moment of decision, e.g., a vote, a deliberation, or an aggregation of preferences, and its single output becomes what the AI does.

How it works

Scroll the diagram sideways to see all of it.

Limitation

Minority Erasure

Aggregating many views into one output generally erases minority positions, since averaging pulls the result toward the middle and views held by few drop out regardless of how strongly they are held. The output also carries no record that anyone disagreed, so users downstream cannot tell that the question was contested at all.

Methods

Handing the decision to a convened public

4 papers

A body of affected people is constituted — an open crowd on a voting platform, a stratified sample, a citizens' assembly, or a panel drawn by lot — and put through a procedure: proposing and voting on statements, or talking them through until a common position forms. Whatever that body settles on becomes the governing input to the model: a constitution to fine-tune on, a set of policy targets, or the pool of raters whose comparisons become the training data. The claim is legitimacy by procedure — the rule binds because of who decided it and how it was decided, not because of what it says.

Built out100%
Adherence38%

Putting the model in the facilitator's seat

6 papers

A model takes over jobs a human facilitator, chair or annotator would do around a group's disagreement. It drafts a candidate group statement from everyone's opinions and critiques and redrafts it as participants rank the versions, picks slates of statements that represent the whole spread of opinion under formal fairness guarantees, chairs a live session by managing the speaking queue and intervening on incivility, voices stakeholders who are not in the room, and rates each contribution for justification, novelty and openness in place of the human coders who used to. Some of this runs live inside a session and some works offline from a collected set of opinions; what the family shares is that the model occupies a role in the deliberation rather than consuming its output afterward.

Built out100%
Adherence38%

Writing a voting rule into the training objective

7 papers

The standard pipeline pools everyone's comparisons into one reward model, which quietly averages disagreeing people into a single stand-in judge. This family makes an explicit collective decision rule the thing the system optimizes instead: maximize the worst-off group's reward, find the policy that beats every rival head to head, draw from a maximal lottery, aggregate whole policies rather than reward numbers, match the population's actual distribution of views, or — in the limiting case, where the object designed is an economic redistribution rule rather than a language policy — optimize directly to win an election among the people the rule governs. The same machinery is turned back on itself to say what any such rule can deliver: the axioms it satisfies, how far it can fall short of the ideal, whether one labeler who misreports can wreck it, and impossibility results showing that no protocol satisfies everyone.

Built out100%
Adherence25%

Voting among many judges at evaluation time

2 papers

Rather than building the aggregation into training, these systems gather many evaluative signals after the fact and combine them with a voting rule instead of an average. Several models from different families, or one compact reward model fitted per annotator, each score the candidate outputs and vote; the same move treats existing benchmark and tournament results as ballots and searches for the ranking of agents that contradicts the fewest of them. Because the voters live at evaluation time, who votes is a knob that can be turned without retraining anything — which is what separates this from rewriting the objective.

Built out100%
Adherence—

Designing what each person is asked

4 papers

Instead of aggregating whatever judgments people happen to volunteer, the system authors the question each person faces, because the question is where the normative work happens. A model runs an open-ended interview, invents the edge case that would be most informative, walks someone through critiquing an agent's behavior until their own preferences take shape, or asks what value they actually brought to a hard case and then which value is wiser here. The designs differ in how the question is chosen, and the difference matters: some pick the next comparison adaptively for information gain, while one applies a single fixed framing — choosing principles without knowing which position you would occupy — that pushes answers toward protecting the worst-off. What separates all of them from voting platforms is that the agenda is written by the system, not by the participants.

Built out100%
Adherence38%

Revising cases and principles against each other

1 paper

Here the collective answer is not a tally at all. Concrete judgments about particular cases and candidate general principles are adjusted against each other until the two cohere — modeled formally under measurable account, systematicity and faithfulness criteria, read as a description of what constitutional AI already does, and shown to be the same search as coherence-driven self-improvement. Attached to the machinery are the arguments for why the target should be justified reasons rather than counts: that preferences do not represent values, and that alignment should be settled by the claims people can fairly make on each other.

Built out100%
Adherence0%

Theoretical foundations

Core concepts

Deliberation, Not Tallying

Jürgen Habermas, Between Facts and Norms (1992)

This concept locates legitimacy in the exchange of reasons among equals: participants justify their views to each other, free of coercion, and remain able to change their minds. Counting votes or averaging preferences does not produce it, because the reasoning is what does the work.

General Will vs. Sum of Preferences

Jean-Jacques Rousseau, The Social Contract (1762)

This distinction separates what is good for a public in common from the sum of what its members privately want. Aggregating private wants, however fairly, does not produce the former, since the two can and often do diverge.

Veil of Ignorance

John Rawls, A Theory of Justice (1971)

This device requires that rules be chosen without knowledge of which position the chooser will occupy under them, e.g., rich or poor, majority or minority. Not knowing generally pushes the chooser to protect the worst-off position, since it may turn out to be their own.

Reasonable Rejection

T. M. Scanlon, What We Owe to Each Other (1998)

This test holds a principle justified only if no individual could reasonably reject it, and it is applied person by person rather than in aggregate. One sufficiently strong objection therefore outweighs many mild preferences, which is the case averaging handles wrongly.

Communicative vs. Strategic Action

Jürgen Habermas

This distinction separates speech aimed at reaching understanding from speech aimed at winning. On Habermas's account only the first confers legitimacy, so a procedure that cannot tell them apart cannot say whether its result was legitimately produced.

Public Reason

Rawls (1993); Habermas

This requirement holds that decisions binding on everyone be justified in terms that citizens holding different comprehensive doctrines can each accept. Reasons internal to one doctrine do not qualify, however sincerely they are held.

The Veil of Ignorance at Runtime

John Rawls (1971)

This applies the veil at inference rather than at design time: the model serves the person in front of it without relying on their position, and acts to their benefit under that ignorance. It is distinct from using the veil once to elicit principles from human participants.

The Difference Principle

John Rawls (1971)

This principle permits inequalities only where they improve the position of the worst-off. Applied to a decision procedure, it makes the worst-off outcome, rather than the average one, the quantity to be maximized.

Overlapping Consensus

John Rawls (1993)

This is agreement on a set of shared norms reached by people who hold incompatible comprehensive doctrines and who endorse those norms for different reasons of their own. The agreement is on the norms, not on the grounds for them.

Reflective Equilibrium

Rawls (1971); coherentist roots in Goodman

This is the state reached by revising particular judgments and general principles against each other until the two are consistent. Neither side is fixed in advance, and the process stops when no further revision is called for.

Coherent Extrapolated Volition

Eliezer Yudkowsky (2004)

This proposal takes the target of alignment to be not what people currently want but what they would want if they were better informed, reasoned more clearly, and had more time to reflect. The extrapolation is meant to be carried out by the system itself.

Consent and the Right to Revoke

John Locke (1689)

This holds that authority is legitimate only where the governed have consented to it, and that the consent may be withdrawn if the authority fails in its trust. The right to revoke is what distinguishes the arrangement from mere subjection.

Source works

Jürgen Habermas

Between Facts and Norms

1992 (Eng. 1996)

Jürgen Habermas

The Theory of Communicative Action

1981

John Rawls

Political Liberalism

1993

Thomas Hobbes

Leviathan

1651

John Locke

Two Treatises of Government

1689

Jean-Jacques Rousseau

The Social Contract

1762

John Rawls

A Theory of Justice

1971

Eliezer Yudkowsky

Coherent Extrapolated Volition

2004

John Rawls

Political Liberalism (overlapping consensus, public reason)

1993

T. M. Scanlon

What We Owe to Each Other (reasonable rejection)

1998

Nelson Goodman

Fact, Fiction, and Forecast (coherentist roots of reflective equilibrium)

1955