SIGN IN SIGN UP

ggml-cuda : use graph allocator (#2684)

use a different function for no_alloc to avoid breaking backwards compat, fixes lora

remove 512 n_batch limit

fixed 2048 batch size

cleanup

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
S
slaren committed
1123f7fbdfb8012e46f05e903e6f675922916378
Parent: ef3f333
Committed by GitHub <noreply@github.com> on 8/22/2023, 1:25:19 PM