SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Kernel][Model] Tune fused_moe Triton configs for Qwen3-30B A3/A3B on H100 (FP8/BF16) (#26268)

Signed-off-by: Shivam <shivampr.dev@gmail.com>
S
shivampr committed
4d0f2661135959775d5b59315134e343d066d3ae
Parent: e93ff6c
Committed by GitHub <noreply@github.com> on 10/20/2025, 2:48:01 PM