SIGN IN SIGN UP

Keep MXFP4 weights quantized on XPU when use_kernels is set (#47923)

The megablocks XPU kernel consumes MXFP4 experts natively, so dequantizing to
bf16 only costs memory and bandwidth.

Co-authored-by: Marc Sun <57196510+SunMarc@users.noreply.github.com>
J
jiqing-feng committed
e15ea9b9e9c3d7692d2f4a2bc606cd9a0776e1af
Parent: d65f37c
Committed by GitHub <noreply@github.com> on 8/26/2026, 5:42:52 PM