SIGN IN SIGN UP

hybrid: recover 6 GiB CUDA OOM and free prefetch on ubatch change

Keep 256 MiB MMQ/FA headroom before allocating expert-prefetch slots.
Return nullptr from CUDA pools on OOM, catch graph-compute bad_alloc,
and fall back from cmoe prefill to decode ubatch instead of aborting.
Release prefetch staging when runtime ubatch changes. Tighten start1660
KVFlash pool and expert-S for the GTX 1660 Ti.

Assisted-by: Grok
A
andi committed
c68217bc6266a034ad690153639306be3af0fd9a
Parent: 2cf3a75