Unlock: AI Control: Safety Without Assuming Alignment
AI control evaluates the whole protocol around a capable model against deliberate attempts to cause a prohibited outcome, instead of assuming the model shares the intended goals. This page draws the trusted boundary, compares protocol components, derives the violation rates of releasing versus deferring flagged work once the audit budget runs out, derives honest-mode task success under deferral, bounds risk across many dependent steps, and lays out what a control evaluation should report.
12 Prerequisites0 Mastered0 Working12 Gaps
Prerequisite mastery0%
Recommended probe
Basic Logic and Proof Techniques is your weakest prerequisite with available questions. You haven't been assessed on this topic yet.
Not assessed18 questions
Not assessed16 questions
Not assessed5 questions
Not assessed30 questions
What Alignment MeansFrontier
Not assessed4 questions
Sign in to track your mastery and see personalized gap analysis.