ggml : same IQ4_NL quantization for CPU/CUDA/Metal (#6196)
* Make quantize_row_iq4_nl do the same thing is quantization on CUDA * Make quantize_row_iq4_nl do the same thing is quantization on CUDA This time for real. backend-ops tests pass. * Now fix test-quantize-fns --------- Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
K
Kawrakow committed
cfd3be76e37dab92c846d75a2421178f20db4a11
Parent: 5b7b0ac
Committed by GitHub <noreply@github.com>
on 3/21/2024, 12:59:38 PM