Taxonomy / Adversarial proceduralism / Debate before an under-informed judge

Debate before an under-informed judge

Two systems are handed opposite sides of a question and argue it out in front of a judge who lacks what they have — the source passage, the evidence, the time, or the raw capability — and so cannot settle the question alone. Because each side gains by exposing the other's errors, the judge hears the strongest case against every claim before deciding, and this line of work measures whether judges actually end up more accurate: on passages the judge has not read, on contested COVID and climate claims before people with prior beliefs, and against a single unopposed consultant. One variant drops fixed sides altogether — the two systems are directed to locate the exact point they disagree on and hand the judge that crux rather than a finished argument.

The method, against Adversarial proceduralism

Scroll the diagram sideways to see all of it.

Concept Analysis: Theoretical Foundations

Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.

Counter-power by design

PartialClaimed · partial

Definition · James Madison, Federalist No. 51 (1788)

Don't rely on written rules or good intentions: build institutions that oppose each other, so that each has both the means and the motive to check the others. "Ambition must be made to counteract ambition."

Analysis

The means half is genuinely built and the motive half is only simulated, which is why this reads PARTIAL rather than PRESENT. What is real: each debater can surface the quote the other suppressed or name the step it skipped, and NEW-S47 isolates that the opposition, not the eloquence, does the work — self-play against a live opponent lifts judge accuracy, persuasion training without an opponent does not. That is a measured mechanism, not a gesture, and it is why this is PARTIAL and not ABSENT. What is missing is the thing Fed. 51 is actually about. Madison's opening premise is that you cannot rely on "parchment barriers" or on the goodwill of whoever administers the system; the remedy is to give the person in the seat "a personal motive to resist encroachments," so that "the interest of the man must be connected with the constitutional rights of the place." The check then operates whether or not anyone running the government wants it to, because the ambition doing the checking belongs to the checker and cannot be withdrawn by the designer. In debate the opposition is entirely installed and entirely revocable by the operator. The two debaters are typically the same model instantiated twice and pointed in opposite directions by an assignment; neither holds a position, an interest, or a stake it brought with it. The approach's own papers demonstrate the revocability: NEW-S47's control condition makes the opposition vanish by removing the opponent, and NEW-S170 keeps the means while instructing the win motive away. A Madisonian check is precisely one that the administrator cannot switch off by reconfiguration. This one is switched off in two of the seven papers, by configuration. So what is built is counter-power by configuration standing where counter-power by design should be — the sides oppose each other only for as long as whoever runs the protocol keeps assigning them to, which returns the arrangement to the good intentions Madison wrote the passage to escape. A second, smaller shortfall in the same direction: Madison's arrangement is that "each may be a check on the other," mutually and without remainder. Here the lateral check between the two advocates is the whole of it; nothing in the structure has means or motive against the seat that actually decides. (The unchecked judge is scored separately as C-44, so I do not rest the downgrade on it — the installed-and-removable motive is sufficient on its own.) Claimed remains YES on the independent axis: Irving asserts that the adversarial structure rather than debater quality is what lets a weak judge supervise strong agents, and the framing borrows Madison's vocabulary directly. The claim is asserted; it is the durability of the motive that is not carried.

Non-domination

AbsentNot claimed · out of scope

Definition · Philip Pettit, Republicanism (1997)

You are unfree if someone holds unchecked power over you, even benevolently. What matters is not whether power is used well but whether the governed can contest it.

Analysis

The only seat holding real power is the judge's, and no protocol here reaches it: the verdict is final, the question and the answer key are fixed outside the proceeding, and nobody bound by the outcome can reopen it. The serious counter-reading is that a weak judge checking strong debaters is Pettit's structure pointed at AI — NEW-S165's consultancy baseline is literally the trust-the-benevolent-advisor condition that debate refuses. But what debate gives the judge is better information, not a standing capacity to contest power: the debaters check each other's claims while holding no power over one another, and NEW-S46 shows the judge's grip disappears once the information asymmetry does, which is precisely the contingency Pettit rules out. No paper frames the work in terms of freedom from arbitrary power; the stated aim throughout is oversight accuracy.

Agonism

PartialNot claimed

Definition · Chantal Mouffe, The Democratic Paradox (2000)

Deep disagreement is permanent and should not be dissolved into consensus. A healthy system converts enemies into adversaries — opponents whose conflict stays alive and legitimate — rather than declaring the argument settled.

Analysis

The carried half is the staged contest under common rules — equal standing for both positions, an opponent rather than an enemy — and NEW-S170 produces the most agonistic artifact in the set: an isolated crux, the unresolved point handed forward in articulated form instead of being dissolved. Everything else runs the opposite way. Positions are assigned rather than held (NEW-S164 sets each expert to argue a different answer, so the disagreement belongs to nobody and dies with the episode), and every paper scores success as judge accuracy against one correct answer, with NEW-S170 naming convergence as the goal and the crux as the residue when convergence fails. Even NEW-S48, which works on genuinely contested COVID and climate claims where judges hold real priors, counts moving skeptical judges toward the mainstream verdict as the win — conflict as a solvent for disagreement, not a permanent condition to be housed.

Incentive compatibility

PartialClaimed · partial

Definition · Leonid Hurwicz; Eric Maskin (mechanism design)

Design the procedure so that honest behavior is each participant's best strategy. The rules of the game do the enforcement, instead of trusting the participants to be good.

Analysis

The design bet is explicitly mechanism-design: structure the game so that the honest side wins because a liar faces an opponent holding the same evidence, letting the rules rather than the debaters' good faith do the enforcing. NEW-S47 supplies the best evidence for that bet — training to win against a live opponent improves judge accuracy while persuasion training without one does not. What is missing is any assurance the equilibrium is actually honest. NEW-S165 prices the leak directly (52% of consultancy errors trace to debaters obfuscating the relevant evidence, and the authors expect this to worsen as debaters improve), NEW-S164 optimizes persuasiveness itself and rests truthfulness on a measured correlation rather than on the rules, and NEW-S170 drops the incentive entirely and instructs models to weigh the evidence candidly — reliance on directed good behavior being the thing mechanism design exists to replace.

Papers

Findings about it

What existing systems were observed to do. These characterise the problem the approach addresses without introducing a mechanism.

On scalable oversight with weak LLMs judging strong LLMs

Zachary Kenton et al., Jul 2024

arXiv:2407.04622101 citationsAnalysisBuilt

Runs debate, consultancy and plain question-answering against each other with weaker models standing in for human judges, across information, mathematics, coding, logic and multimodal asymmetries: debate beats consultancy on every task, and beats direct question-answering only where the judge lacks information the debaters have.

AI Debate Aids Assessment of Controversial Claims

Salman Rahman et al., Jun 2025

arXiv:2506.02175AnalysisBuilt

Puts two AI systems to debate opposing sides of contested COVID-19 and climate claims in front of human judges holding mainstream or skeptical prior beliefs, and finds debate beats a single AI consultant by 4–10% on accuracy — up to +15.2% for mainstream judges, and +4.7% even for skeptical judges who initially had the claim wrong.

Debating with More Persuasive LLMs Leads to More Truthful Answers

Akbir Khan et al., Feb 2024

arXiv:2402.06782AnalysisBuilt

Two stronger models that hold the information needed to answer each argue for a different answer while a weaker model or a human who lacks that information picks between them, with debate reaching 76% accuracy for the model judges and 88% for the human judges against naive baselines of 48% and 60%.

Debate Helps Supervise Unreliable Experts

Julian Michael et al., Nov 2023

arXiv:2311.08702AnalysisBuilt

Collects human-written debates on hard reading comprehension questions where the judge has not read the source passage and sees only the arguments and the short quotes the debaters selectively reveal, and finds judges reach 84% accuracy under debate against 74% under a single consultant, with debates running 68% of the length.