SIGN IN SIGN UP

perf: shrink prefill graphs and retune start1660 for 1660 Ti

Grow/shrink scheduler compute buffers on cmoe phase switch so a 1856
prefill ubatch no longer keeps its VRAM peak through decode. Evict and
recapture CUDA graphs on topology change instead of disabling them after
instantiate OOM. Skip STARTED leftover-draft prefill so turn-2 stays on
the decode graph.

DFlash standalone inject now requires consecutive draft positions, and
unset --spec-type is inferred from the draft GGUF. start1660 uses S=28,
prefill 1856, decode 64, and n_max=2 from the measured A/B.
A
andi committed
972f7d9efa91c775ce000fea9f71ad0fe95428b0
Parent: a79242e