5 · Serve and the GPU lie · 29/38
vLLM: batch, pages, TTFT
KV pool occupancy is not nvidia-smi VRAM.
playReserved = one max-size block. Paged = grab pages on demand. More requests fit.allocationReserved mode: every request checks out a full max-size block, most of it never used. Flip to paged and watch how many more requests fit.
KV in usereserved · wastedfree page
—requests fit
—memory wasted
—KV used / allocated
dashboard sees · pages allocated—
pool checked out — looks full in both modes
actually working · KV in use—
pages holding real tokens
Reserved blocks waste pages. PagedAttention packs the pool.