Legible deliberation
Aristotelian practice, made inspectableWhen virtues conflict, the weighing between them is done in the open — you can see which considerations were weighed and why one prevailed, rather than just receiving the verdict.
Taxonomy / Phronetic adjudication
When the AI's own values conflict, its trained practical judgment arbitrates case by case, and the deliberation is left visible rather than hidden.
Scroll the diagram sideways to see all of it.
Limitation
An AI that justifies choices by appeal to its own practical wisdom can dress any output in a plausible-sounding deliberation after the fact. There is no independent standard to check the reasoning against, so a genuinely wise judgment and a confabulated one look identical from the outside.
Instead of collapsing a trade-off into a single number before the question is ever asked, these systems produce, for the particular situation in front of them, an explicit account of which values bear on it and which way each one cuts. Kaleido stops at the list, naming the values, rights and duties in play and whether each supports or opposes the action; chain-of-thought training on value conflicts carries the same material through to a verdict; the automated parliament splits the work across several delegates who each speak for a moral perspective, rewrite one another's answers, and then settle on one together. Because the considerations are set out rather than hidden in a score, they can be read back at scale: sorting hundreds of trolley-problem traces by the kind of reason given shows models reasoning from duties in the chain of thought while explaining themselves by consequences afterward.
Rather than committing to one ethical theory, the system holds several and settles their disagreements with a rule written down in advance: split your confidence across the theories and let those weights shape what a learning agent is rewarded for, or treat the theories as parties bargaining over lotteries and take the bargaining solution. Only the theories' verdicts on the case are inputs; the rule that combines them does not change from case to case. That is the deliberate opposite of open-ended judgment, and the point of it — the answer is predictable and can be argued about before any particular case arises.
These papers argue that what is missing is not a procedure to bolt on at answer time but a trait the model does not have, and that training should aim at the trait itself. One holds that fine-tuning installs surface behavior — refuse this, comply with that — and never the ability to reason when those instructions collide, which is exactly the seam adversarial prompts pry open. The other holds that the harder gap is knowing which approach a problem with no settled method calls for, and sketches how that kind of judgment might be benchmarked, trained and built. They name different traits, and neither has been built, but both make the same move: change what training installs rather than the shape of any one answer.
When virtues conflict, the weighing between them is done in the open — you can see which considerations were weighed and why one prevailed, rather than just receiving the verdict.
Right action lies between too much and too little — courage sits between recklessness and cowardice — and where the mean falls depends on the case. Finding it is a judgment made per situation, not a fixed rule.