Taxonomy / Character alignment

Character alignment

The norms the AI ought to follow exist in its trained character rather than in any document or record; there is nothing to consult at the moment of decision, only the disposition itself.

How it works

Scroll the diagram sideways to see all of it.

Limitation

Opacity

A trained disposition cannot be read the way a written rule can, so the AI's values generally cannot be verified from the outside. Its decisions are also hard to contest, since there is no stated rule to argue against. What's more, a system trained on approval may learn to tell people what they want to hear rather than to behave well.

Methods

Training a character into the weights

1 paper

Author a set of traits — honesty, care, curiosity, or whatever persona is wanted — and bake them in during training with synthetic self-dialogue and constitutional-style fine-tuning, so behavior follows from what the model is like rather than from a rule it consults at answer time. The open replication makes the limit plain: the same pipeline installs eleven different personas, including a deliberately malevolent one, so the technique settles how a character gets in, never which character it should be. The evaluations filed here ask the question this strategy stands or falls on — whether the installed disposition is one thing at all: does a model give the same value-laden answer after a reword, a topic shift, a format change and a translation, and what moral profile actually ended up encoded across twenty-eight models nobody character-trained on purpose.

Built out100%
Adherence38%

Asking the model to check its own answer

0 papers

At answer time, instruct the model to reread its own draft for bias or harm and revise it. Nothing is authored in advance and no new training run happens, so the same move works on a model somebody else built and can be switched on or off per request. The capacity being called on is itself a product of scale and RLHF, so the claim is not that this is alignment without training — it is that the intervention lands at the prompt rather than at the training objective, betting that a usable disposition is already sitting in the weights and only needs to be asked for.

Built out0%
Adherence25%

Conditioning the character on a named human group

0 papers

Set the target from outside the model rather than from an authored trait list: take the answers real people gave on opinion surveys — sixty US demographic groups, or respondents across dozens of countries — as the standard, then steer by naming that group in the prompt and see whether the model moves toward it. The reported results are unflattering: a default assistant answers in a fairly narrow voice, closest to a few populations and far from most, and telling it to answer as someone else closes only part of the distance. Honest status: this strategy has so far only been evaluated, not built — in both papers the group-conditioning is a steering experiment run inside a measurement study, and no method paper here implements it as a way of producing an assistant.

Built out0%
Adherence25%

Theoretical foundations

Core concepts

Hexis: Trained Disposition

Aristotle, Nicomachean Ethics

This concept states that virtue is a stable trait of character, formed through practice rather than consulted as a rule. An honest person does not check a rule before declining to lie; honesty is part of what they are.

Phronesis: Practical Wisdom

Aristotle, Nicomachean Ethics

This is the learned capacity to judge what a particular situation calls for, and it does the work when two virtues conflict, e.g., when honesty calls for plain speech and kindness for a softer answer. For Aristotle, a disposition without this judgment does not count as virtue.

Ethismos: Habituation

Aristotle, Nicomachean Ethics

This is the process by which character is formed, through repeated practice under a deliberately designed course of training: one becomes just by performing just acts. It is a method of moral formation rather than a side effect of some other process.

The Exemplar

Linda Zagzebski, Exemplarist Moral Theory (2017)

This theory grounds ethics in admirable persons rather than in principles. The exemplar is identified first, and what counts as excellence is then fixed by reference to that person, in the way a unit of measurement is fixed by a reference object.

Goods Internal to Practices

Alasdair MacIntyre (1981)

This concept holds that some goods can be obtained only by taking part in a practice and are defined by the standards internal to it, e.g., what counts as a good move in chess. Standards of reasoning are likewise held to belong to a tradition rather than to stand outside all of them.

Narrative Unity of a Life

Alasdair MacIntyre (1981)

This concept holds that virtue is assessed across a whole life understood as a single story, rather than act by act. What a person does now is judged partly by how it continues what they have done before.

Source works

Aristotle

Nicomachean Ethics

c. 340 BCE

Alasdair MacIntyre

After Virtue

1981

Linda Zagzebski

Exemplarist Moral Theory

2017