CUDA: optimize MMQ int8 tensor core performance (#8062)
* CUDA: optimize MMQ int8 tensor core performance * only a single get_mma_tile_x_k function * simplify code, make functions constexpr
J
Johannes Gäßler committed
9a590c82262dd518137f85406e65e452fdf2aca3
Parent: 52fc870
Committed by GitHub <noreply@github.com>
on 6/24/2024, 10:41:23 AM