llama : improve infill support and special token detection (#9798)
* llama : improve infill support ggml-ci * llama : add more FIM token strings ggml-ci * server : update prompt on slot restore (#9800) * gguf : deprecate old FIM token KVs
G
Georgi Gerganov committed
11ac9800aff532715a5bc7991062c68ba3472e6e
Parent: 943d20b
Committed by GitHub <noreply@github.com>
on 10/12/2024, 5:21:51 AM