Paper2021
Learning Transferable Visual Models From Natural Language Supervision
Radford et al.
Trains an image encoder and a text encoder together to predict which captions match which images across 400 million pairs, producing representations (CLIP) that transfer to new visual tasks with no task-specific training.
48 pageslink checked 17 Sept 2026FreeAdvanced