Paper2017
Proximal Policy Optimization Algorithms
Schulman et al.
Replaces trust-region policy optimisation's constrained second-order update with a clipped surrogate objective that a handful of stochastic gradient steps can optimise, trading some theoretical guarantee for practical simplicity.
link checked 17 Sept 2026FreeIntermediate