SIGN IN SIGN UP

Custom RoPE + bettter memory management for CUDA (#2295)

* Custom RoPE + bettter memory management for CUDA

* Adjusted look ahead in ggml_cuda_pool_malloc to 5%

This is sufficient it seems.
We end up using about 200 MB less VRAM that way when running
the 13B model with context 8192.

---------

Co-authored-by: Iwan Kawrakow <iwan.kawrakow@gmail.com>
K
Kawrakow committed
d924522a46c5ef097af4a88087d91673e8e87e4d
Parent: 4d76a5f
Committed by GitHub <noreply@github.com> on 7/21/2023, 2:27:51 PM