feat: allow KVFlash paging on MLA K-only caches
Ling/bailingmoe3 stores compressed MLA K and reconstructs V, so the pager no longer requires a V tensor on every attention layer.
A
andi committed
cd88aad7870697f64a4f544fa78205ccc04b2beb
Parent: 972f7d9