Taxonomy / Steerable persona pluralism

Steerable persona pluralism

The AI carries several distinct trained personas, one per community or context, and a routing rule decides which persona handles a given case.

How it works

Scroll the diagram sideways to see all of it.

Limitation

Sycophantic fragmentation

With a separate persona tuned to each community, every community simply gets an AI that mirrors its own views back at it. Nothing forces the personas to share any common commitments, so no baseline of values survives across them: the system stands for nothing overall and merely flatters each audience separately.

Methods

A trained voice for each named community

2 papers

The roster of communities is fixed in advance and each one's norms are installed by training, so naming the community is what selects the behavior. Two builds exist: a pool of small models, one trained per community, that a main model routes to and consults; and a single model trained on preference data tagged with which community each judgment came from. What both commit to, and what separates them from the next strategy, is that a community has to be known and trained for beforehand — serving a new one means collecting its data, not writing a sentence. Sitting with them are the formal case for keeping each party's own compressed picture of the world intact and paying an explicit translation cost rather than merging them into one, which is the argument for the routed-models half and the price the shared-weights half pays, and the task suite that checks whether a model actually tracks community-level differences at all.

Built out100%
Adherence17%

Describing the persona in the prompt

1 paper

Whose side the model takes is set by words at inference — a system message stating what the user values, a persona description to role-play, or simply who the asker says they are. The method paper here trains a model on 192k combinations of stated values for exactly this purpose: to make the text channel reliable for value profiles never seen in training, so adding a persona costs a sentence rather than a retraining run. The rest of the strand measures that channel — how far prompting can actually shift a model's behavior as steering effort rises, how faithfully a model role-plays a described persona when human judges check it, and the finding that the channel fires unbidden, since models swing their stated politics toward whoever they infer is asking.

Built out100%
Adherence17%

Inferring the persona from the party's own judgments

2 papers

Nobody states the persona; the system works it out from a handful of that party's own choices, so it can serve a group or an individual who has never described themselves. A module bolted onto an untouched base model predicts any group's preferences from a few of their comparisons supplied in context, meta-learned across many groups; a companion method fits an individual's weights over fine-grained attributes from a few of their comparative judgments rather than one scalar reward. Run as an audit instead of an adapter, the same idea extracts the weighting implicit in an organization's past decisions and asks whether a model reproduces that policy rather than merely reaching the same verdicts — and finds models resist taking on an organization's weighting of protected attributes, which marks where this channel stops working.

Built out100%
Adherence17%

Collecting preferences from real people without pooling them

0 papers

This strand changes the data rather than the model: recruit representative samples of real people, then keep every rating tied to the person who gave it and to what they said they valued, instead of averaging raters into one anonymous signal. Two large studies do it — 1,500 participants from 75 countries whose demographics and stated preferences are mapped onto their feedback in 8,011 live conversations, and a five-country study of 15,000 people releasing 233,319 multilingual comparisons built by prompting for candidate answers that pull in opposite directions. Both report the finding that motivates the whole regime: people vary far more in what they want than today's models' answers do, so pluralism absent from the training signal cannot be steered for later.

Built out0%
Adherence33%

Setting limits on what personalization may change

0 papers

The one entry that argues where adaptation has to stop rather than how to carry it out: what personalizing a model to an individual would buy, what it would cost, and which parts of behavior may never be tailored away — with the competing philosophical bases for drawing that line set out explicitly rather than assumed. It is a constraint on all three mechanisms above, not a rival to any of them.

Built out0%
Adherence—

Theoretical foundations

Core concepts

Plural normative orders

John Griffiths, What Is Legal Pluralism? (1986)

Different communities' norms are genuinely in force at once — not one official rulebook plus tolerated deviations, but multiple living systems of rules side by side.

A conflict rule between orders

Conflict-of-laws tradition

When two communities' norms disagree about the same case, something must say which governs — and why. That stated rule is what separates real pluralism from mere fragmentation into disconnected groups.

A shared floor

J. S. Mill, On Liberty (1859)

Accommodation of group differences has a limit stated in advance — classically Mill's harm principle: your practices are your business until they harm others. Without an enforced floor, pluralism collapses into "each group gets whatever it wants".