Keep MXFP4 weights quantized on XPU when use_kernels is set (#47923)
The megablocks XPU kernel consumes MXFP4 experts natively, so dequantizing to bf16 only costs memory and bandwidth. Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>
J
jiqing-feng committed
e15ea9b9e9c3d7692d2f4a2bc606cd9a0776e1af
Parent: d65f37c
Committed by GitHub <noreply@github.com>
on 8/26/2026, 5:42:52 PM