SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Bugfix] fixes the decoding metadata of dense mla's fp8 kvcache. (#27144)

Signed-off-by: Tao He <linzhu.ht@alibaba-inc.com>
Signed-off-by: Lucas Wilkinson <lwilkins@redhat.com>
Co-authored-by: Lucas Wilkinson <lwilkins@redhat.com>
T
Tao He committed
250fb1b8ea836ab4cdd2890581815a9c19b8b8ed
Parent: 647214f
Committed by GitHub <noreply@github.com> on 10/21/2025, 6:27:03 PM