Taxonomy / Deliberative aggregation / Designing what each person is asked

Designing what each person is asked

Instead of aggregating whatever judgments people happen to volunteer, the system authors the question each person faces, because the question is where the normative work happens. A model runs an open-ended interview, invents the edge case that would be most informative, walks someone through critiquing an agent's behavior until their own preferences take shape, or asks what value they actually brought to a hard case and then which value is wiser here. The designs differ in how the question is chosen, and the difference matters: some pick the next comparison adaptively for information gain, while one applies a single fixed framing — choosing principles without knowing which position you would occupy — that pushes answers toward protecting the worst-off. What separates all of them from voting platforms is that the agenda is written by the system, not by the participants.

The method, against Deliberative aggregation

Scroll the diagram sideways to see all of it.

Concept Analysis: Theoretical Foundations

Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.

Deliberation, Not Tallying

PartialClaimed · partial

Definition · Jürgen Habermas, Between Facts and Norms (1992)

This concept locates legitimacy in the exchange of reasons among equals: participants justify their views to each other, free of coercion, and remain able to change their minds. Counting votes or averaging preferences does not produce it, because the reasoning is what does the work.

Analysis

Two of the five designs get people to reason rather than report a want: the moral-graph interview asks which of two considerations is wiser in a given context, and the reflective-dialogue method has users critique agent behavior until their own view shifts, so reasons are elicited and minds do change. But no participant ever faces another participant — every exchange is between one person and a model that sets the agenda, so no view is ever justified to the person it would burden, and each design terminates in a count or a fit (wisdom edges settled by vote, rewards by regression). GATE and the active-query learner elicit no reasons at all, only a specification and a scalar. Claimed is IMPLIED because the moral-graph paper advertises a legitimate alignment target grounded in participants' own judgments rather than counted preferences, a benefit that holds only if reasoning is doing the work.

General Will vs. Sum of Preferences

PartialClaimed · partial

Definition · Jean-Jacques Rousseau, The Social Contract (1762)

This distinction separates what is good for a public in common from the sum of what its members privately want. Aggregating private wants, however fairly, does not produce the former, since the two can and often do diverge.

Analysis

Two designs aim at an object that is not a sum of wants — principles chosen without knowing one's own position, and judgments of which consideration is wiser rather than which is preferred — and the moral-graph method builds a single shared graph whose winners govern everyone. The other three aim the opposite way by design: GATE and the active-query learner exist to pin down one user's private specification, and the reflective-dialogue method deliberately ends with a separate reward model per user and no procedure for bringing them into contact. Even where a common object is reached, it is frozen at survey time or settled by tallying pairwise wisdom votes, so nothing is maintained as a standing public judgment. The framing borrows the vocabulary — values rather than preferences, fairness chosen from behind a veil — without any source asserting the distinction outright.

Veil of Ignorance

PartialClaimed · partial

Definition · John Rawls, A Theory of Justice (1971)

This device requires that rules be chosen without knowledge of which position the chooser will occupy under them, e.g., rich or poor, majority or minority. Not knowing generally pushes the chooser to protect the worst-off position, since it may turn out to be their own.

Analysis

One design runs the device literally: roughly 2,000 participants choose governing principles for an assistant without knowing which party they would be under them, and their choices shift toward protecting the worst-off, exactly the effect Rawls predicts. The other four run the inverse on purpose — values cards are drawn from situations the participant was actually in, GATE fits the specification to this user's own position, and the active learner selects comparisons precisely for what they reveal about this particular person's reward. And even in the veil study the screen covers human participants at design time only; the deployed model applies the resulting principles while knowing exactly whose request it is handling. Claimed is YES because that paper's own title and framing assert the device as its method.

Reasonable Rejection

AbsentNot claimed · out of scope

Definition · T. M. Scanlon, What We Owe to Each Other (1998)

This test holds a principle justified only if no individual could reasonably reject it, and it is applied person by person rather than in aggregate. One sufficiently strong objection therefore outweighs many mild preferences, which is the case averaging handles wrongly.

Analysis

No design in this set puts a rule to a person and asks whether they could reject it. The veil study yields worst-off-protective principles, but that is a convergent outcome of a different device rather than a rejectability test, and the moral graph's surviving cards look unbeaten only because pairwise wisdom judgments are counted — a card a minority holds intensely loses to many mild contrary judgments, which is precisely the aggregation Scanlon's test exists to block. The single-user methods have nobody at the table for a specification to be justified to, and the active-query learner explicitly resolves every answer into one scalar that is traded off against the rest of the fit. No source invokes the test or advertises a benefit that would require it.

Concept Analysis: Newly Introduced

Values elicited as what was attended to

Added

Instead of asking what people want or what principles they endorse, the interview asks what they actually paid attention to when they had to decide. That produces a unit — the values card — that is neither a preference nor a principle, and the pairwise wisdom judgments over those cards are a form of collective reasoning no deliberative tradition specifies.

Questions chosen for what they reveal

Added

The tradition's participants arrive with their views already formed, and where it thinks about questions at all it thinks about who sets the agenda. Neither Habermas nor Rawls has anything about optimizing the question itself. Sadigh's volume-removal objective synthesizes the comparison expected to be most informative about the reward, and GATE has the model invent the edge case a user would not have thought to raise. In both, the design of what to ask is the thing being optimized.

Papers