k_quants tuning for Falcon-7b (#2816)
* Make ggml-cuda.cu build with QK_K = 64 Using LLAMA_CUDA_FORCE_DMMV = ON and -nommq it runs and produces a meaningful result. * k_quants tuning for Falcon-7b --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
K
Kawrakow committed
a6d1189fdd4c1ab4ba23f9d777f8950901dcffb2
Parent: c48c5bb
Committed by GitHub <noreply@github.com>
on 8/27/2023, 12:19:59 PM