CPU/CUDA: Gemma 2 FlashAttention support (#8542)
* CPU/CUDA: Gemma 2 FlashAttention support * apply logit_softcap to scale in kernel * disable logit softcapping tests on Metal * remove metal check
J
Johannes Gäßler committed
e11bd856d538e44d24d8cad4b0381fba0984d162
Parent: 8f824ff
Committed by GitHub <noreply@github.com>
on 8/24/2024, 7:34:59 PM