SIGN IN SIGN UP

llama : support requantizing models instead of only allowing quantization from 16/32bit (#1691)

* Add support for quantizing already quantized models

* Threaded dequantizing and f16 to f32 conversion

* Clean up thread blocks with spares calculation a bit

* Use std::runtime_error exceptions.
K
Kerfuffle committed
4f0154b0bad775ac4651bf73b5c216eb43c45cdc
Parent: ef3171d
Committed by GitHub <noreply@github.com> on 6/10/2023, 7:59:17 AM