Artificial Intelligence
Learning by acting, with delayed and sparse feedback.
10 topics · 1 curated work
Primary sources and full treatments.
The standard text: builds the whole field from the bandit problem up to function approximation and policy gradients.