Project Sherlock

Artificial Intelligence · AI Alignment & Safety

Corrigibility

A topic within AI Alignment & Safety, itself one of 11 topics in that field and part of Artificial Intelligence.

Reading on Corrigibility

3

A way in

  1. Start here

    No prior grounding assumed.

    AI 'Stop Button' Problem

    Computerphile · 2017

    Explains, for a general audience, why building a reliable off-switch into a goal-directed AI system is harder than it sounds, because a sufficiently…

  2. Go deeper

    Primary sources and full treatments.

    Corrigibility

    Soares et al. · 2015

    Defines a corrigible agent as one that does not resist being shut down, corrected, or modified, and shows that naive utility-maximising designs fail…

    +1 more at this level

3 works

Video2017

AI 'Stop Button' Problem

Computerphile

Explains, for a general audience, why building a reliable off-switch into a goal-directed AI system is harder than it sounds, because a sufficiently capable agent has an instrumental incentive to prevent itself being switched off.

13 minuteslink checked 17 Sept 2026
Paper2015

Corrigibility

Soares et al.

Defines a corrigible agent as one that does not resist being shut down, corrected, or modified, and shows that naive utility-maximising designs fail to have this property even when explicitly given a shutdown button.

10 pageslink checked 17 Sept 2026
Paper2016

The Off-Switch Game

Hadfield-Menell, Dragan, Abbeel & Russell

Models the shutdown problem as a game between a human and a robot, and shows a robot that is uncertain about the human's true objective and treats the human's decision to switch it off as evidence of that objective will let itself be switched off.

7 pageslink checked 17 Sept 2026

Other topics in AI Alignment & Safety