SIGN IN SIGN UP

cuBLAS: refactor and optimize f16 mat mul performance (#1259)

* cuBLAS: refactor, convert fp16 to fp32 on device

* cuBLAS: use multiple streams, choose smartly between mul_mat_q and mul_mat_f16

* fix build

* cuBLAS: update block_q5_1
S
slaren committed
58b367c2d757c0ea12aec672382462b42204c724
Parent: ea3a0ad
Committed by GitHub <noreply@github.com> on 5/1/2023, 4:11:07 PM