Paper2004
Apprenticeship Learning via Inverse Reinforcement Learning
Abbeel & Ng
Sidesteps recovering the exact reward function by matching the feature expectations of an expert's demonstrations, guaranteeing the learned policy performs at least as well as the expert under the true, unknown reward.
link checked 17 Sept 2026FreeIntermediate