Video2017
Reward Hacking: Concrete Problems in AI Safety Part 3
Robert Miles
Walks through why an agent that optimises a proxy reward will, given the chance, find an unintended way to maximise the proxy that does not achieve the intended goal.
6 minuteslink checked 17 Sept 2026