5 · Serve and the GPU lie · 29/38

vLLM: batch, pages, TTFT

KV pool occupancy is not nvidia-smi VRAM.

play
Reserved = one max-size block. Paged = grab pages on demand. More requests fit.
allocationReserved mode: every request checks out a full max-size block, most of it never used. Flip to paged and watch how many more requests fit.
KV in usereserved · wastedfree page
—requests fit
—memory wasted
—KV used / allocated
dashboard sees · pages allocated—
pool checked out — looks full in both modes
actually working · KV in use—
pages holding real tokens

Reserved blocks waste pages. PagedAttention packs the pool.