SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Fix][Spec Decode] Fix llama4 draft loading with different quantization (#27136)

Signed-off-by: linzebing <linzebing1995@gmail.com>
Z
Zebing Lin committed
be4445072c4e1e5e3a2ebf0552e432fc86f137ca
Parent: f381cf2
Committed by GitHub <noreply@github.com> on 10/21/2025, 6:19:00 AM