3 · Train a model you own · 23/38

nanoGPT

The training loop and estimate_mfu.

play
one step · MFU = compute / wall · tiny batch looks “busy” and still wastes the GPU
MFU — · smi —

estimate_mfu in nanoGPT is this bar. Point at it before talking util.

800 steps MPS: val BPB 5.21→3.18. Prefill 78k tok/s vs decode 683 (114×). MFU vs A100 = 0.37%.