Paper2010
Understanding the Difficulty of Training Deep Feedforward Neural Networks
Glorot & Bengio
Traces the difficulty of training deep networks to how weight initialisation interacts with activation functions and layer-to-layer signal variance, and derives an initialisation scheme (Xavier initialisation) that keeps that variance stable across depth.
8 pageslink checked 17 Sept 2026FreeAdvanced