SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Bugfix] Disable FlexAttention direct block mask building for encoder-only models (#27344)

Signed-off-by: Isotr0py <mozf@mail2.sysu.edu.cn>
I
Isotr0py committed
084a9dae801c6756cb50c03daa313bca96fa660b
Parent: c9461e0
Committed by GitHub <noreply@github.com> on 10/22/2025, 4:39:08 PM