Project Sherlock

Artificial Intelligence

AI Alignment & Safety

Whether these systems do what we intend — the field's most consequential open problem.

11 topics · 20 curated works

Topics

Reading in AI Alignment & Safety

20

A way in

  1. Start here

    No prior grounding assumed.

    Reward Hacking: Concrete Problems in AI Safety Part 3

    Robert Miles · 2017

    Walks through why an agent that optimises a proxy reward will, given the chance, find an unintended way to maximise the proxy that does not achieve…

    +2 more at this level

  2. Then

    Assumes you know the vocabulary.

    Concrete Problems in AI Safety

    Amodei et al. · 2016

    Reframes AI safety as five tractable engineering problems rather than a speculative long-term concern.

    +5 more at this level

  3. Go deeper

    Primary sources and full treatments.

    The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents

    Nick Bostrom · 2012

    Argues intelligence and final goals are independent (the orthogonality thesis), and that agents with almost any final goal will converge on similar…

    +10 more at this level

12 of 20 works

Video2017

AI 'Stop Button' Problem

Computerphile

Explains, for a general audience, why building a reliable off-switch into a goal-directed AI system is harder than it sounds, because a sufficiently capable agent has an instrumental incentive to prevent itself being switched off.

13 minuteslink checked 17 Sept 2026
Paper2017

AI Safety Gridworlds

Leike et al.

Builds a suite of small gridworld environments, each designed so an agent that maximises reward the naive way visibly fails a safety property, turning specification gaming into something you can measure rather than just discuss.

13 pageslink checked 17 Sept 2026
Paper2012

The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents

Nick Bostrom

Argues intelligence and final goals are independent (the orthogonality thesis), and that agents with almost any final goal will converge on similar instrumental subgoals — self-preservation, resource acquisition, goal-content integrity — which is what makes a highly capable system's specific goal so consequential.

15 pageslink checked 17 Sept 2026
Paper2015

Corrigibility

Soares et al.

Defines a corrigible agent as one that does not resist being shut down, corrected, or modified, and shows that naive utility-maximising designs fail to have this property even when explicitly given a shutdown button.

10 pageslink checked 17 Sept 2026

Also covered elsewhere

This subject genuinely sits in more than one domain. These fields approach the same ground with different methods.

Elsewhere in Artificial Intelligence