SIGN IN SIGN UP

cuBLAS: use host pinned memory and dequantize while copying (#1207)

* cuBLAS: dequantize simultaneously while copying memory

* cuBLAS: use host pinned memory

* cuBLAS: improve ggml_compute_forward_mul_mat_f16_f32 with pinned memory

* cuBLAS: also pin kv cache

* fix rebase
S
slaren committed
7fc50c051ae8a78e9643fdf172d12e20f2dd9b6c
Parent: b1ee8f5
Committed by GitHub <noreply@github.com> on 4/29/2023, 12:04:18 AM