3 · Train a model you own · 25/38

Scaling laws

Slide C. ★ is the knee. Right of ★ = undertrained.

play
iso-compute · ★ is Chinchilla-optimal · right of ★ = too big, undertrained
N — · D — · tok/p — · L —

C ≈ 6ND. Right of ★ = GPT-3’s mistake: too big, too few tokens.

C ≈ 6ND. GPT-3 sat right of this knee.