6 · Giant MoE (8×H100) · 34/38
Pipeline bubble & all-reduce
Idle at 100%. The wire is the step.
playmicrobatchesstages
nvidia-smi says—
a stage with any resident kernel reads busy — even mid-bubble
pipeline efficiency · goodput—
filled cells ÷ total = M / (M + S − 1)
The step is the slowest GPU plus the wire. 100% util can be waiting on NCCL.