PROMPT, TOKENIZED
hey
my
friend
x = [ 3 tokens × 3584 ]
3 tokens, each one an integer
looked up in the embedding table
this matrix is what enters layer 1
hi
generated — now fed back in alone
LAYER 1 — one of 28, each with its own weights
x
W_Q
W_K
W_V
Q 3 × 3584
K 3 × 512
V 3 × 512
one row
per token
attention scores,
then discarded
kept
KV CACHE
hey
my
friend
L1
K
V
L2
⋮
L28
28 layers
× tokens
× K and V
hi
2 (K and V) × 512 × 28 layers × 2 bytes (FP16) = 57,344 bytes per token
≈ 56 KB, held in GPU memory until the request finishes
← Prev
Next →