Paper2020
Learning to Summarize from Human Feedback
Stiennon et al.
Applies reward modelling and reinforcement learning from human comparisons to summarisation, and finds optimising directly for human preference beats optimising for the ROUGE metric summarisation had been judged on for years.
link checked 17 Sept 2026FreeIntermediate