SIGN IN SIGN UP

Vulkan k-quant mmq and ggml-backend offload functionality (#6155)

* Fix Vulkan no kv offload incoherence

* Add k-quant mul mat mat shaders

* Rework working buffer allocation, reduces vram use noticeably

Clean up cpu assist code, replaced with ggml-backend offload function

* Default to all dedicated GPUs

* Add fallback for integrated GPUs if no dedicated GPUs are found

* Add debug info which device is allocating memory

* Fix Intel dequant issue

Fix validation issue

* Fix Vulkan GGML_OP_GET_ROWS implementation

* Clean up merge artifacts

* Remove Vulkan warning
0
0cc4m committed
ba0c7c70ab5b15f1f2be7fb0dfbe0366dda30d6c
Parent: d48ccf3
Committed by GitHub <noreply@github.com> on 3/29/2024, 4:29:21 PM