SIGN IN SIGN UP

Use fused D256 attention for Qwen prefill tails

Route eligible full-precision Qwen 3.5-family prefill tails through MLX
0.32.2's fused head-dimension-256 kernel. This keeps off-boundary chunk
attention bit-identical to one-shot attention while reducing tail latency
and transient allocation. Decode, array masks, short initial prompts, and
quantized cache implementations retain their existing paths.
J
Jinho Jang committed
cc085536e8d47efa0934377a472411b4fbe72484
Parent: 45e9a6a