Paper2018
GLUE: A Multi-Task Benchmark and Analysis Platform for Natural Language Understanding
Alex Wang et al.
Argues that no single task measures general language understanding and assembles nine tasks into one benchmark with a shared leaderboard, which then became the standard a model had to clear before claiming general competence.
link checked 17 Sept 2026FreeIntermediate