SIGN IN SIGN UP

Porting the improved K-Quant CUDA kernels to OpenCL (#1966)

* Added broken new q4k quant

* xx + ib0

* Fix q2_k fast kernel

* Use preprocessor for QK_K

* Add q6_k fast matmul kernel

* ported q3k speedup successfully

* ported q2k and q5k speedups

* remove old dot kernels and template

* fixed global const struct types

* fixing address spaces

* fixed string too long CI issue

---------

Co-authored-by: 0cc4m <picard12@live.de>
L
LostRuins committed
96a712ca1b7f427e3bd7ffc0c70b2105cfc7fbf1
Parent: d3494bb
Committed by GitHub <noreply@github.com> on 6/29/2023, 3:56:43 AM