D · Deep nets (refresh) · 13/38

Sequences and vanishing gradients

g₀/g_T → 0 unless a gate holds ≈ 1.

play
gradient back through time · RNN dies · LSTM gate holds ≈ 1
g₀ / g_T —

Product of |w|<1 vanishes. The gate is why LSTMs lasted.

Depth in time kills the gradient. Attention is one hop.