SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[NVIDIA] [Perf] Update to leverage flashinfer trtllm FP4 MOE throughput kernel (#26714)

Signed-off-by: jiahanc <173873397+jiahanc@users.noreply.github.com>
Co-authored-by: Michael Goin <mgoin64@gmail.com>
J
jiahanc committed
41d3071918bd6dcb6259bac8a617cfed45d42707
Parent: fb5e10d
Committed by GitHub <noreply@github.com> on 10/16/2025, 11:20:25 PM