cuBLAS: use host pinned memory and dequantize while copying (#1207)
* cuBLAS: dequantize simultaneously while copying memory * cuBLAS: use host pinned memory * cuBLAS: improve ggml_compute_forward_mul_mat_f16_f32 with pinned memory * cuBLAS: also pin kv cache * fix rebase
S
slaren committed
7fc50c051ae8a78e9643fdf172d12e20f2dd9b6c
Parent: b1ee8f5
Committed by GitHub <noreply@github.com>
on 4/29/2023, 12:04:18 AM