Taxonomy / Deliberative aggregation / Voting among many judges at evaluation time

Voting among many judges at evaluation time

Rather than building the aggregation into training, these systems gather many evaluative signals after the fact and combine them with a voting rule instead of an average. Several models from different families, or one compact reward model fitted per annotator, each score the candidate outputs and vote; the same move treats existing benchmark and tournament results as ballots and searches for the ranking of agents that contradicts the fewest of them. Because the voters live at evaluation time, who votes is a knob that can be turned without retraining anything — which is what separates this from rewriting the objective.

The method, against Deliberative aggregation

Scroll the diagram sideways to see all of it.

Papers