SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Performance] Dual stream execution of "shared_experts" and "selected_experts" inside FusedMoE (#26440)

Signed-off-by: Alexander Matveev <amatveev@redhat.com>
A
Alexander Matveev committed
344a0017c06e40041efe35ccff7670c1131067e3
Parent: becb7de
Committed by GitHub <noreply@github.com> on 10/21/2025, 9:38:29 PM