1 · Architectures — common core, different bets · 18/38
GQA and MLA — shrinking KV
Decode is a memory problem. Fewer KV, or a latent.
playmodel: 32 layers · 32 query heads · head-dim 128 · fp16 · KV budget 40 GB
GQA shares K/V. MLA stores a latent. Long context is bought here.