llama : fix embeddings (#5796)
* llama : fix embeddings ggml-ci * llama : do not use KV cache for non-causal models ggml-ci * embeddings : fix llama_batch_init arg * llama : add pooling switch * llama : distinguish token vs sequence embeddings ggml-ci * llama : assert pooling tensor * llama : simplify causal mask condition ggml-ci * llama : assert input batch with pooling enabled * readme : update API changes list
G
Georgi Gerganov committed
29ae62d2ae163e2b68aa0ad3bf2ab4636de0c957
Parent: e0843af
Committed by GitHub <noreply@github.com>
on 3/4/2024, 8:31:20 PM