SIGN IN SIGN UP

cuda : performance optimizations (#1530)

* xor hack

* block y dim

* loop unrolling

* Fixed cmake LLAMA_CUDA_BY option

* Removed hipblas compatibility code

* Define GGML_CUDA_DMMV_BLOCK_Y if not defined

* Fewer iters, more ops per iter

* Renamed DMMV X/Y compilation options
J
Johannes Gäßler committed
1fcdcc28b119a6608774d52de905931bd5f8a43d
Parent: ac7876a
Committed by GitHub <noreply@github.com> on 5/25/2023, 9:07:29 PM