SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[cpu] Dispatch un-quantized linear to oneDNN/ACL by default for AArch64 (#27183)

Signed-off-by: Fadi Arafeh <fadi.arafeh@arm.com>
Co-authored-by: Michael Yang <Michael.Yang@arm.com>
F
Fadi Arafeh committed
163965d183a980f28909003fac32100d53e1d2e9
Parent: a03cf9b
Committed by GitHub <noreply@github.com> on 10/21/2025, 2:02:58 AM