SIGN IN SIGN UP

CUDA: optimize MMQ int8 tensor core performance (#8062)

* CUDA: optimize MMQ int8 tensor core performance

* only a single get_mma_tile_x_k function

* simplify code, make functions constexpr
J
Johannes Gäßler committed
9a590c82262dd518137f85406e65e452fdf2aca3
Parent: 52fc870
Committed by GitHub <noreply@github.com> on 6/24/2024, 10:41:23 AM