SIGN IN SIGN UP

ggml : add IQ2 to test-backend-ops + refactoring (#4990)

* ggml : add IQ2 to test-backend-ops + refactoring

ggml-ci

* cuda : update supports_op for IQ2

ggml-ci

* ci : enable LLAMA_CUBLAS=1 for CUDA nodes

ggml-ci

* cuda : fix out-of-bounds-access in `mul_mat_vec_q`

ggml-ci

* tests : avoid creating RNGs for each Q tensor

ggml-ci

* tests : avoid creating RNGs for each tensor

ggml-ci
G
Georgi Gerganov committed
38566680cdfe982a495562332c25b9227de9cf8d
Parent: ba69bbc
Committed by GitHub <noreply@github.com> on 1/17/2024, 4:54:56 PM