TP>1: stop pinning the single-card KV_MEM, warn on DFLASH_TOKENS>7 (#40)
choki-lin measured the pin stranding ~16 GiB across 2x3090 (137,210 tokens of pool vs 302,223 from GPU_UTIL, all decode deltas inside spread) and a -27% C1 regression at DFLASH_TOKENS=15/TP=2 with the lookup tail accepting nothing. The launcher now sizes from GPU_UTIL under tensor parallelism unless the user pins, and warns on long verify blocks at TP>1 until the tail regression is diagnosed. README multi-GPU section rewritten around their controlled A/B; dual-GPU pointers updated.
M
mhenrichsen committed
60921ede72ba214c3c62ffab7317834346718a2a
Parent: 04edd2b