Vanishing Gradient
/ˈvænɪʃɪŋ ˈɡreɪdiənt/
A problem in deep networks where gradients become extremely small, preventing effective learning in early layers.
In italiano: Vanishing GradientVanishing gradients occur when repeated multiplication of small derivatives (< 1) makes gradients exponentially smaller in backpropagation. Common in deep RNNs and networks with sigmoid activations. Solutions: ReLU, LSTM, ResNet.
Examples
- Deep RNNs failing to learn long dependencies
- Sigmoid activation in deep networks
- Pre-ResNet very deep networks