Project Sherlock

Artificial Intelligence · Deep Learning

Scaling Laws

A topic within Deep Learning, itself one of 14 topics in that field and part of Artificial Intelligence.

Reading on Scaling Laws

2

2 works

Paper2020

Scaling Laws for Neural Language Models

Kaplan et al.

Shows language model loss follows smooth power laws in model size, dataset size, and compute across seven orders of magnitude, and that within a fixed compute budget it is better to train large models on comparatively modest data and stop before convergence.

30 pageslink checked 17 Sept 2026
Paper2022

Training Compute-Optimal Large Language Models

Hoffmann et al.

Retrains the earlier Kaplan scaling laws with a wider sweep of model and data sizes and finds most large language models of the time were significantly undertrained relative to their size — model size and training tokens should scale together, roughly equally, as compute grows.

36 pageslink checked 17 Sept 2026

Other topics in Deep Learning