THE OLD WAY — reserve the worst case up front
used
reserved for a 32,768-token context that will never arrive — wasted
one request, 500 tokens used, 1.8 GB reserved → ~98% of it idle
PHYSICAL GPU MEMORY — fixed 16-token blocks
REQUEST A — logical view
0–15
16–31
32–47
token positions — each block holds 16 tokens' worth of K and V
contiguous from 0, which is a convenient fiction
REQUEST B — logical view
0–15
16–31
BLOCK TABLE A
logical → physical
0–15 → block 7
16–31 → block 23
32–47 → block 4
BLOCK TABLE B
0–15 → block 12
16–31 → block 15
BLOCK TABLE B
0–15 → block 7 ← shared
16–31 → block 15
same system prompt → same blocks, refcount 2
HOST (CPU) MEMORY
ordinary system RAM — blocks copied
off the GPU and parked here over PCIe,
~50× slower than the GPU's own memory
copied out to
CPU memory
15
B's private block
swapped back in —
to a different block
empty again
BLOCK TABLE B — rewritten
0–15 → block 7 still shared
16–31 → block 2 ← moved
same logical view, new physical home
← Prev
Next →