SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Bugfix][Qwen] fixes the weights dtype in qwen3_next: it is actually a bfloat16 (#27030)

Signed-off-by: Tao He <linzhu.ht@alibaba-inc.com>
T
Tao He committed
bde9e2272a28342d30fe9c4de4c1c9633ff153d8
Parent: 0840560
Committed by GitHub <noreply@github.com> on 10/17/2025, 3:37:52 AM