SIGN IN SIGN UP

ggml-cuda : perform cublas mat mul of quantized types as f16 (#3412)

* ggml-cuda : perform cublas matrix multiplication of quantized types as fp16

* rename CC_TURING to CC_VOLTA

* disable fp16 mat mul completely with multi GPU
S
slaren committed
f5ef5cfb18148131fcf45bdd2331f0db5ab7c3d0
Parent: 40e07a6
Committed by GitHub <noreply@github.com> on 9/30/2023, 4:12:57 PM