SIGN IN SIGN UP

quantize : add '--keep-split' to quantize model into shards (#6688)

* Implement '--keep-split' to quantize model into several shards

* Add test script

* Update examples/quantize/quantize.cpp

Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>

* Split model correctly even if tensor id is out-of-order

* Update llama_model_quantize_params

* Fix preci failures

---------

Co-authored-by: z5269887 <z5269887@unsw.edu.au>
Co-authored-by: Georgi Gerganov <ggerganov@gmail.com>
J
jiez committed
1966eb2615242f224bf9ca939db8905ab6a174a0
Parent: 784e11d
Committed by GitHub <noreply@github.com> on 4/25/2024, 10:29:35 AM