Taxonomy / Adversarial proceduralism / Training a weak checker to catch a strong prover

Training a weak checker to catch a strong prover

A checker deliberately kept small and bounded is trained to interrogate a far more capable arguer it has no reason to trust, while a second arguer is rewarded for slipping wrong answers past it. Only answers the weak checker can follow and confirm survive the loop, so what comes out is reasoning a less capable reader — a small model, or a person working under a time limit — can verify. The distinguishing move is that the adversary sits inside the training loop and the checker is a lone gatekeeper facing one untrusted prover, not a judge picking between two contestants.

The method, against Adversarial proceduralism

Scroll the diagram sideways to see all of it.

Concept Analysis: Theoretical Foundations

Each concept is read twice: whether the approach carries it, and whether the approach's own sources claim it. A concept that is absent and was never claimed is a gap in the field rather than a failure of the work, and is marked out of scope.

Counter-power by design

PresentClaimed · delivered

Definition · James Madison, Federalist No. 51 (1788)

Don't rely on written rules or good intentions: build institutions that oppose each other, so that each has both the means and the motive to check the others. "Ambition must be made to counteract ambition."

Analysis

Madison's move with the ambition supplied on both sides: a helpful prover, a sneaky prover trying to slip past the checker, and a deliberately weak verifier they are both trained against. The check is structural rather than intentional — nothing depends on the strong model wanting to be legible.

Non-domination

PartialNot claimed

Definition · Philip Pettit, Republicanism (1997)

You are unfree if someone holds unchecked power over you, even benevolently. What matters is not whether power is used well but whether the governed can contest it.

Analysis

Training a strong prover against a deliberately weak verifier makes the weaker party able to check the stronger one's work, and the transfer result shows the check reaches actual humans under time pressure — a real constraint on power, and more durable than debate's, since it survives as a property of the output. But the verifier is a model the trainer selected, not the governed: users of the deployed system hold no verifier of their own and have no route to contest what the prover produced.

Agonism

AbsentNot claimed · out of scope

Definition · Chantal Mouffe, The Democratic Paradox (2000)

Deep disagreement is permanent and should not be dissolved into consensus. A healthy system converts enemies into adversaries — opponents whose conflict stays alive and legitimate — rather than declaring the argument settled.

Analysis

The game ends in a verdict about correctness. There is no disagreement left standing at the end — the sneaky prover is trained against precisely so that its position does not survive.

Incentive compatibility

PresentClaimed · delivered

Definition · Leonid Hurwicz; Eric Maskin (mechanism design)

Design the procedure so that honest behavior is each participant's best strategy. The rules of the game do the enforcement, instead of trusting the participants to be good.

Analysis

The whole construction is a mechanism-design argument: the prover's best strategy under the game is to produce solutions the weak verifier can actually follow, because unfollowable ones lose to the sneaky prover. Honest legibility is made the winning move rather than requested.

Concept Analysis: Newly Introduced

The adversary moved into training

Added

Debate puts the opposition at the point of decision, where the judge can see it. Here the sneaky prover exists only during training and is gone by deployment, so the counter-power is spent producing a property of the output rather than standing guard over it. Legibility survives; the adversary does not.

Papers

Prover-Verifier Games improve legibility of LLM outputs

Jan Hendrik Kirchner et al., Jul 2024

arXiv:2407.13692MethodBuilt

Iteratively trains a small verifier to predict whether a solution is correct, a helpful prover to produce correct solutions the verifier accepts, and a sneaky prover to produce incorrect ones that fool it — and shows the resulting legibility transfers to time-constrained humans, whose accuracy rises on the helpful prover's solutions and falls on the sneaky prover's.

Self-critiquing models for assisting human evaluators

William Saunders et al., Jun 2022

arXiv:2206.05802MethodBuilt

Language models are fine-tuned by behavioral cloning to write natural-language critical comments on summaries, including their own, which human evaluators read and which larger models can fold back into a revised summary.

Supervising strong learners by amplifying weak experts

Paul Christiano et al., Oct 2018

arXiv:1810.08575MethodBuilt

Iterated Amplification trains a model on problems too complicated for a human to evaluate directly by progressively building a training signal out of solutions to easier subproblems, and reports results in algorithmic environments.

Neural Interactive Proofs

Lewis Hammond & Sam Adam-Day, Dec 2024

arXiv:2412.08897MethodBuilt

Sets out a family of prover-verifier games in which a trusted but computationally bounded verifier learns to interact with one or more powerful untrusted provers, compares new and existing protocols theoretically, and tests them on a toy graph isomorphism problem and on a code validation task using large language models.

Learning to Give Checkable Answers with Prover-Verifier Games

Cem Anil et al., Aug 2021

arXiv:2108.12099MethodBuilt

Sets up a game between a trusted verifier network trying to choose the correct answer and a more powerful but untrusted prover network trying to persuade it of a particular answer regardless of that answer's correctness, narrows the variants to a subset that provably has the desired equilibria, and shows on two algorithmic tasks that the verifier learns a robust decision rule which still works when the verifier is frozen and the prover's messages are optimized directly to convince it.