llama : support requantizing models instead of only allowing quantization from 16/32bit (#1691)
* Add support for quantizing already quantized models * Threaded dequantizing and f16 to f32 conversion * Clean up thread blocks with spares calculation a bit * Use std::runtime_error exceptions.
K
Kerfuffle committed
4f0154b0bad775ac4651bf73b5c216eb43c45cdc
Parent: ef3171d
Committed by GitHub <noreply@github.com>
on 6/10/2023, 7:59:17 AM