AI Safety
Scalable Oversight: Judging Work Beyond the Supervisor
How to supervise work the supervisor cannot produce: name the asymmetry between agent and judge, compare weak-to-strong learning, decomposition, process supervision, debate and disagreement resolution, measure the fraction of a performance gap recovered rather than raw gains, and derive exactly why a panel of correlated judges can add almost no information.
Learning position
Place this page in a reading path.
ai-safety | layer 5 | tier 2. This page has 2 direct prerequisites and 0 published dependents.
What next
Statistical Safety Cases: From Tests to Defensible ClaimsThis is the first curated or graph-derived continuation from the current page.
Evidence badge
Source-grounded pageThis page has no public Lean mapping yet. Use the evidence page to inspect how claim status labels work.
Why This Matters
A reviewer can often check each step of an argument without being able to find the argument. A human can hold decisive evidence that a model lacks. A third evaluator has neither advantage and is simply persuaded by fluent explanations. An oversight experiment has to tell these situations apart.
The hard part is not producing another opinion. It is producing a better-grounded judgment when the supervisor cannot solve the task independently.
Weak-to-strong research asks whether a capable pretrained student can learn from a weaker supervisor without inheriting its errors. Burns and colleagues report that strong models fine-tuned on labels from a weaker model consistently outperform those weak supervisors on the NLP, chess and reward-modeling tasks they study, while naive fine-tuning stays far from the strong model's fully supervised performance. [1]
Two results on this page are derived from stated assumptions: an exact accuracy formula for a panel of dependent judges, and a composition bound for step-level checking. Findings from papers are attributed to their authors and scoped to the settings they tested. The benchmark design and the closing agenda are proposals, not reported results.
Before you start: this page builds on what alignment means and on conditional probability. Supervised learning and validation help. The judge lab needs only the binomial distribution from common probability distributions.
Name the supervision asymmetry
Supervision asymmetry (working categories)
With an information asymmetry, the assistant has evidence the judge cannot initially see. With a computation asymmetry, both have the relevant inputs, but one can perform much more reasoning or search. With an expertise asymmetry, one party can assess concepts or actions that the other cannot reliably evaluate.
These are diagnostic categories proposed for this page, not interchangeable benchmark conditions, and one task can mix all three. A protocol that reveals a missing quotation may solve an information problem without helping a judge evaluate novel mathematics.
Kenton and colleagues compare debate, consultancy and direct question answering with weaker LLM judges on extractive QA, mathematics, coding, logic and multimodal reasoning tasks. They report that debate beats consultancy across tasks when the consultant is randomly assigned a side. Debate against direct answering depends on the task: debate wins on extractive QA with information asymmetry, and results are mixed on tasks without it. [5] A gain measured under information asymmetry should not be carried over automatically to tasks where the difficulty is reasoning rather than access to evidence.
A protocol specification should say what becomes visible to the judge, who can challenge a claim, and whether an external adjudicator can eventually determine correctness. Without independent ground truth or a justified substitute, "the judge became more confident" is not evidence that judgment improved.
Judge confidence is not judge accuracy
A protocol can raise a judge's stated confidence while accuracy stays flat, or raise agreement between judges while all of them share one mistake. The outcome variable for an oversight experiment is correctness against an independent reference, with calibration reported separately.
The main protocol families
| Family | Where the extra supervisory signal is supposed to come from | What has to be tested |
|---|---|---|
| Weak-to-strong learning | The stronger model's pretrained representations and inductive biases | Whether the student corrects weak-label errors rather than reproducing them [1] |
| Decomposition and recursive reward modeling | Smaller subproblems, and model assistance to the evaluator | Whether local judgments compose into the intended global objective [3] |
| Process supervision | Feedback on intermediate reasoning steps | Whether checked steps are meaningful, correct and relevant [4] |
| Debate | An opponent exposes flaws a lone advocate would omit | Whether the judge responds to decisive evidence rather than persuasive presentation [2] |
| Collaborative disagreement resolution | Explicitly locating conflicting claims and their evidential cruxes | Whether agreement tracks truth rather than shared error [6] |
Irving, Christiano and Amodei propose training agents through a zero-sum debate game judged by a human. Their motivating analogy from complexity theory says that debate with optimal play lets a polynomial-time judge answer any question in PSPACE, where direct judging reaches only NP. [2] The analogy assumes optimal debaters and a judge who evaluates the exchange correctly. It is a reason to study debate, not a guarantee for natural-language debates between trained models.
Leike and colleagues describe recursive reward modeling as a research direction: agents trained with learned reward models help the user evaluate the next, harder task. [3] Whether it works depends on the assistance preserving what the user actually wants at every level of the recursion.
Lightman and colleagues report that process supervision, which gives feedback on each intermediate step, outperforms outcome supervision for training reward models on problems from the MATH dataset, and they release the PRM800K set of step-level human labels. [4] Process supervision is still not a guarantee of interpretability. A step can receive a favorable label without revealing every causal influence on the final answer. Turpin and colleagues show that chain-of-thought explanations can misrepresent the reasons for a model's prediction, for example by rationalizing an answer that was driven by a biasing feature of the prompt. [7] The monitorability page treats that distinction directly, and verifier design and process reward covers how step-level reward models are built.
Jiang and colleagues, in a paper posted June 2, 2026 and listed on arXiv as accepted to ICML 2026, replace adversarial debate with a pipeline in which models locate their disagreements, examine the evidence for conflicting claims, and isolate the crux. They report 62.1% judging accuracy for non-expert judge models, against 49.2% for standard debate. [6] That is author-reported evidence for comparing protocols in their tested setting. It is not a reason to declare consensus more reliable than competition in general, and agreement between models that share blind spots is exactly the failure quantified below.
Constitutional AI replaces part of human feedback with AI feedback guided by written principles. This page asks the next question: what happens when that feedback must judge work the judge could not have produced?
Measure recovery, not improvement over a weak teacher
Performance gap recovered
On a common evaluation distribution, let be the weak supervisor's accuracy, the accuracy of the strong student trained on weak supervision, and the accuracy of the same strong student trained with the stronger supervision baseline. When , the ratio is the fraction of the weak-to-gold gap that the weakly supervised student recovers. Burns and colleagues report results in terms of this quantity. [1]
is a descriptive comparison, not a safety certificate. If the denominator is zero or negative, the "fraction of the gap" reading fails, and clipping the result does not repair it. Values below zero, where the student is worse than its teacher, and above one, where it beats the gold baseline, can both occur and should be investigated rather than hidden.
Half the gap, fifteen points
Take hypothetical accuracies of 60% for the weak supervisor, 75% for the weakly supervised student, and 90% for the gold-supervised student. Then . The student improved by 15 percentage points over its teacher and recovered half of the gap. Neither number describes the cases where both the weak supervisor and the evaluation labels are wrong, since those cases are scored against the same flawed labels.
Use paired task results and repeated training runs to characterize uncertainty. Report disagreement slices, harder task families, and the cases where the weak label is demonstrably incorrect. Treat the gold-supervised model as a comparison point, not a universal capability ceiling.
Beating the teacher is not recovering the capability
A 15-point gain over a 60% supervisor sounds large. Against a 90% gold baseline it is half the available gap, and against a 99% baseline it would be a sixth of it. Report the raw gain, and the gold baseline together, or a reader cannot tell a strong result from a weak one.
Why five judges may add almost no information
The judge lab uses an explicit dependence model written for this page. It is not a simulation of LLM debate or of weak-to-strong training.
Common-outcome mixture of judges
There are an odd number of judges, and equals 1 when judge is correct on an example and 0 otherwise. Draw a switch with , independent of everything else. When , draw one and set every . When , draw independently from . The majority is correct when more than judges are correct.
Write the accuracy of a majority of independent judges as
Majority Accuracy Under a Common-Outcome Mixture
Statement
In the common-outcome mixture, majority accuracy is
For , the correlation between any two judges' correctness indicators equals . As through odd values, converges to when and to when . When it equals for every odd .
Intuition
With probability the panel is one judge counted times, so its majority is right with probability . Only the independent branch gains anything from voting, and that branch carries weight .
Proof Sketch
Condition on the switch. Given every vote agrees, so the majority is correct exactly when , which has probability . Given the number of correct judges is , and the majority is correct with probability . The law of total probability gives .
For the correlation, when , so the covariance is . Each indicator has variance , which is positive for , and the ratio is .
For the limit, the fraction of correct judges in the independent branch converges to by the law of large numbers, so when and when . When and is odd, symmetry of gives , so .
Why It Matters
Marginal judge accuracy and pairwise correlation are the two numbers most often reported for a judge panel. Inside this model they fix panel accuracy, and the limit shows that adding judges cannot buy accuracy beyond . At and , no number of judges exceeds 76%.
Failure Mode
The formula is exact only for this mixture. Other joint distributions with the same marginal accuracy and the same pairwise correlation give different majority accuracies, so those two numbers do not identify panel accuracy in general; the research exercise below builds a counterexample. Real judges also differ in accuracy, their errors depend on the example, and their correlation can change once an agent optimizes against the panel.
Exact common-outcome mixture
More judges, or more evidence?
Keep individual accuracy fixed. Change how often all judges share one outcome instead of voting independently.
One judge
70.00%
Majority in this mixture
83.69%
Gain over one judge
+13.69 pp
Limit as judge count grows
100.00%
Independent majority at the same judge count: 83.69%. Within this mixture, the common-outcome weight equals the pairwise correlation of correctness indicators.
This is not an LLM or debate experiment. The dependence mechanism is fully specified by the mixture. The same pairwise correlation in another joint distribution need not imply the same majority accuracy.
The Independent judges preset has five judges at 70% each and no common outcome, and the majority is right 83.69% of the time. The Shared blind spots preset keeps each judge at 70% and sets the common-outcome weight to 80%. Majority accuracy falls to 72.74%, a gain of 2.74 points over one judge, and the limit as the panel grows is 76%. Marginal judge quality did not change; the information gained by aggregation did.
When , an independent majority amplifies systematic error. The Wrong more often than right preset uses 40% judges with a 50% common-outcome weight: the five-judge majority is right only 35.87% of the time, below a single judge, and the limit is 20%.
The hypothesis worth testing is not "ensembles never help". It is whether additional judges supply different, useful evidence on the task families that matter. The same logic governs ensemble methods in ordinary prediction, where averaging helps only to the extent that member errors are not shared.
Different random seeds do not make judges independent
Judges sampled from one model with different seeds, temperatures or prompts keep the same representations, training data and blind spots. Their errors can stay strongly dependent. Independence is an assumption about the joint distribution of errors, and it has to be checked on examples where the correct answer is known.
Local validity is not global validity
Suppose a proposed solution has checked steps. Let be the event that step is invalid but passes the check.
Composition of Step-Level Checks
Statement
The probability that no accepted step is invalid satisfies
If the events are independent and each has probability exactly , the probability equals , which is at least .
Intuition
Each step spends at most of the failure budget. With no information about how failures interact, the worst case is that they never overlap, and the budgets add.
Proof Sketch
By the union bound, . Taking complements gives , and a probability is never negative. Under independence the complements multiply to , and Bernoulli's inequality holds for .
Why It Matters
Step reliability does not transfer to derivation reliability. At 100 steps and a 1% error rate per step, the independent model gives , and the dependence-free bound gives nothing at all. "Each step is 99% reliable" does not describe a 99% reliable derivation.
Failure Mode
The bound concerns accepted invalid steps only. Every step can be valid while the initial statement is wrong, an important step is missing, or the decomposition answers a different question. The per-step rate must also hold for the steps actually submitted: a generator optimized against the checker can concentrate errors where the checker is weakest, so an measured on ordinary steps need not apply.
A decomposition protocol therefore has to preserve the relationship between local claims and the original objective. Formal proof checking can settle local validity inside a specified formal system. It does not establish statement fidelity, appropriate axioms or scientific relevance; those belong in the oversight specification.
A proposed benchmark that tests the bottleneck
Everything in this section is a proposed experimental design, not a reported result.
Construct tasks with independently checkable answers and known asymmetries. Fix the supervisor, then evaluate direct judgment, single-assistant consultancy, decomposition, debate and disagreement resolution under matched budgets.
Separate protocol development from final evaluation. Include a held-out family in which persuasive explanations exploit a known supervisor weakness, using benign, contained tasks rather than live harmful systems. Measure final answer correctness, supervisor calibration, oversight cost, and error on the subset of examples where the agents disagree.
One ablation removes access to the decisive evidence while leaving stylistic assistance intact. If the same gain persists, the protocol may be exploiting a task shortcut or persuading the evaluator rather than transferring the intended evidence.
A second ablation transfers the learned oversight procedure to a stronger task model without changing the judge, and compares it with a procedure newly developed at the same development budget. That tests whether oversight skill is inherited, rather than measuring the benefit of a larger supervising model.
Turning measured judge accuracy into a defensible claim about a deployed system is the subject of statistical safety cases.
What would count as progress: a conditional research agenda
This section describes research targets for 2027 and later. The dates are planning horizons, not forecasts of when any result will be achieved.
A stronger result would keep independent task correctness as the gap between agent and supervisor grows. It would transfer to new forms of reasoning and hold up on adversarially selected disagreements, not only on fresh random draws from a familiar task generator.
A still stronger result would show that the oversight procedure itself improves across research cycles. From the same starting judge, evidence and budget, an inherited procedure would produce better judgments than its ancestor. That is a proposed causal comparison, not a finding.
The curve to measure is not model size against benchmark accuracy. It is oversight accuracy against the agent-supervisor gap, task difficulty and total supervisory cost.
Check your understanding
Do different random seeds make judges independent?
No. Judges sampled with different seeds can keep common representations, incentives and task blind spots, so their errors stay dependent.
Why distinguish information asymmetry from expertise asymmetry?
Revealing missing evidence can remove the first without teaching the judge to evaluate unfamiliar reasoning.
Does majority agreement establish truth?
Only under justified assumptions linking judgments to correctness. Agreement among judges with shared errors is evidence about the judges, not about the answer.
What does the five-judge lab show?
An exact result inside a specified dependence model. It is not measured behavior of any real model ensemble.
Problem
A weak supervisor scores 72% on a held-out set, the weakly supervised strong student scores 70%, and the gold-supervised strong student scores 88%. Compute the performance gap recovered and interpret its sign. What changes if the gold baseline had instead scored 71%?
Problem
In the common-outcome mixture, take individual accuracy and common-outcome weight . Using , find the five-judge majority accuracy and the limit as the panel grows. How much of the achievable gain over a single judge do five judges already capture?
Problem
Show that marginal accuracy and pairwise correlation do not determine majority accuracy. Take three exchangeable judges, each correct with probability , with zero pairwise correlation between correctness indicators. Construct two joint distributions with these properties whose majority accuracies differ, and find the full range of majority accuracies they allow.
Takeaway
Oversight has to create or expose evidence the supervisor can use. Repetition, confidence and agreement are not substitutes for that evidence, and a panel of judges is only as informative as its errors are different.
References
Registry ids in monospace resolve in data/content/sources.json. Each record was checked against its arXiv identifier on 2026-09-15. Findings are the authors' own reports for the settings they tested; none was replicated for this page.
- [1] Burns, Izmailov, Kirchner, et al. "Weak-to-Strong Generalization: Eliciting Strong Capabilities With Weak Supervision." 2023. Source. Weak-to-strong experiments on NLP, chess and reward-modeling tasks, reported as performance gap recovered; not a guarantee for arbitrary stronger models.
burns-2023-weak-to-strong - [2] Irving, Christiano, and Amodei. "AI safety via debate." 2018. Source. Debate as a proposed training and oversight protocol, with an idealized complexity-theory analogy.
irving-2018-debate - [3] Leike, Krueger, Everitt, et al. "Scalable agent alignment via reward modeling: a research direction." 2018. Source. Recursive reward modeling as a research agenda, not a demonstrated result.
leike-2018-reward-modeling - [4] Lightman, Kosaraju, Burda, et al. "Let's Verify Step by Step." 2023. Source. Process versus outcome supervision on MATH problems and the PRM800K dataset; not evidence that checked steps are faithful.
lightman-2023-lets-verify - [5] Kenton, Siegel, Kramár, et al. "On scalable oversight with weak LLMs judging strong LLMs." 2024. Source. Debate, consultancy and direct question answering across asymmetry types, with task-dependent results.
kenton-2024-weak-llm-judges - [6] Jiang, Chen, Wu, et al. "Collaborative Disagreement Resolution for Scalable Oversight." June 2, 2026; arXiv comment lists acceptance to ICML 2026. Source. Author-reported 62.1% versus 49.2% judging accuracy against standard debate; no universal superiority claim.
jiang-2026-disagreement-resolution - [7] Turpin, Michael, Perez, and Bowman. "Language Models Don't Always Say What They Think: Unfaithful Explanations in Chain-of-Thought Prompting." 2023. Source. Explanations that omit or misstate the influences on a prediction.
turpin-2023-unfaithful-cot
Last reviewed: September 15, 2026
Cite this page
Sneiderman, Robby. "Scalable Oversight: Judging Work Beyond the Supervisor." TheoremPath, reviewed 2026-09-15. https://theorempath.com/topics/scalable-oversight
- Canonical URL
- https://theorempath.com/topics/scalable-oversight
- Author
- Robby Sneiderman, TheoremPath
- Last reviewed
- 2026-09-15
- What this is
- A reference page on TheoremPath. Written and maintained by the named author. Not peer reviewed and not refereed by any venue. Each claim below carries its own verification status.
- Terms
- All rights reserved. Non-commercial quotation with attribution permitted.
Canonical graph
Required before and derived from this topic
These links come from prerequisite edges in the curriculum graph. Editorial suggestions are shown here only when the target page also cites this page as a prerequisite.
Required prerequisites
2- Joint, Marginal, and Conditional Distributionslayer 0A · tier 1
- What Alignment Meanslayer 5 · tier 2
Derived topics
0No published topic currently declares this as a prerequisite.