6 · Giant MoE (8×H100) · 34/38

Pipeline bubble & all-reduce

Idle at 100%. The wire is the step.

play
microbatchesstages
nvidia-smi says—
a stage with any resident kernel reads busy — even mid-bubble
pipeline efficiency · goodput—
filled cells ÷ total = M / (M + S − 1)

The step is the slowest GPU plus the wire. 100% util can be waiting on NCCL.