SIGN IN SIGN UP

TP>1: stop pinning the single-card KV_MEM, warn on DFLASH_TOKENS>7 (#40)

choki-lin measured the pin stranding ~16 GiB across 2x3090 (137,210 tokens
of pool vs 302,223 from GPU_UTIL, all decode deltas inside spread) and a
-27% C1 regression at DFLASH_TOKENS=15/TP=2 with the lookup tail accepting
nothing. The launcher now sizes from GPU_UTIL under tensor parallelism
unless the user pins, and warns on long verify blocks at TP>1 until the
tail regression is diagnosed. README multi-GPU section rewritten around
their controlled A/B; dual-GPU pointers updated.
M
mhenrichsen committed
60921ede72ba214c3c62ffab7317834346718a2a
Parent: 04edd2b