feat: enable Ling phase-batching without cpu-moe
All-GPU MoE can use LLAMA_CMOE_PREFILL/DECODE without parking experts on the CPU. start-ling-tiny.sh ships the measured 2048/64 + KVFlash 8192 recipe for the 1660 Ti.
A
andi committed
00dd107ba97a787497f127394a772b7a87cb0fea
Parent: 8cae34d