SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[torch.compile] Enable silu_mul_fp8_quant fusion without custom ops enabled (#27146)

Signed-off-by: zjy0516 <riverclouds.zhu@qq.com>
J
Jiangyun Zhu committed
ab3e80042eac24dd362408e6d63ad98768046359
Parent: ceacedc
Committed by GitHub <noreply@github.com> on 10/22/2025, 4:22:39 AM