1 · Architectures — common core, different bets · 18/38

GQA and MLA — shrinking KV

Decode is a memory problem. Fewer KV, or a latent.

play
model: 32 layers · 32 query heads · head-dim 128 · fp16 · KV budget 40 GB
KV cache per token

GQA shares K/V. MLA stores a latent. Long context is bought here.