D · Deep nets (refresh) · 12/38

Convolutions

Hover a pixel. One kernel, every place.

play
one 3×3 kernel · hover the image · that patch becomes one feature-map pixel
—

Weight sharing. Attention later replaces this for sequences.

Not the LLM path. Here so attention is a solution, not a fashion.