SIGN IN SIGN UP

ggml : update softmax n_task calculation (#5126)

updated the n_task calculation to use max number of
threads possible. This has improved the prompt eval
performance by around 5% for DOT kernels and by
around 10% for MMLA kernels on AWS Graviton3.
S
snadampal committed
7032f4f6349c17a8352f9f93f7d2122f45469e59
Parent: 5f1925a
Committed by GitHub <noreply@github.com> on 1/26/2024, 5:17:59 PM