Paper2015
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals & Jeff Dean
Trains a small model to match the soft output probabilities of a large trained model rather than its hard labels, showing that the large model's mistakes carry information a small model can learn from directly.
link checked 17 Sept 2026FreeIntermediate