SIGN IN SIGN UP

cuda: add q8_0->f32 cpy operation (#9571)

llama: enable K-shift for quantized KV cache
It will fail on unsupported backends or quant types.
I
Ivan committed
116efee0eef09d8c3c4c60b52fa01b56ddeb432c
Parent: 0b3bf96
Committed by GitHub <noreply@github.com> on 9/24/2024, 12:14:24 AM