LinkedIn analytics tracking pixel for AIDOLS AI consulting website performance measurement
Governance & Risk

Scalable Oversight

Scalable oversight is the alignment problem of supervising AI systems on tasks where humans cannot easily check the answer themselves — and the family of techniques being developed to keep oversight tractable as models grow more capable.

Full definition

Approaches include: (1) AI-assisted critique — a second AI flags errors a human can verify; (2) Debate — two AIs argue and a human judges; (3) Recursive reward modeling — break the task into checkable sub-tasks; (4) Weak-to-strong generalization (OpenAI, 2023) — small supervisors successfully aligning larger models; (5) Constitutional AI / RLAIF. The "sandwiching" methodology (Cotra) measures whether non-experts using AI assistance can supervise expert-level tasks.

Why it matters

As enterprises deploy AI in domains where employees cannot easily verify outputs (medical coding, legal contract review, multi-step financial analysis), scalable oversight stops being academic. It is the core question behind every "human in the loop, but the human can't actually check" workflow.

Example

A radiology AI flags potential findings; a junior radiologist with AI critique tools achieves the same review quality as a senior radiologist alone — sandwiching working in practice.

Source & further reading

Primary source: Bowman et al. — "Measuring Progress on Scalable Oversight for Large Language Models" (Anthropic) (2022).

Citation policy: this entry is part of the AIDOLS AI Implementation Glossary and may be quoted for research, journalism, and education with attribution to aidolsgroup.com/de/glossary/scalable-oversight/.