Vanishing Gradient

/ˈvænɪʃɪŋ ˈɡreɪdiənt/

A problem in deep networks where gradients become extremely small, preventing effective learning in early layers.

In italiano: Vanishing Gradient

Vanishing gradients occur when repeated multiplication of small derivatives (< 1) makes gradients exponentially smaller in backpropagation. Common in deep RNNs and networks with sigmoid activations. Solutions: ReLU, LSTM, ResNet.

Examples

  • Deep RNNs failing to learn long dependencies
  • Sigmoid activation in deep networks
  • Pre-ResNet very deep networks