SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

Bugfix - pass 'max_num_tokens_padded' into 'moe_lora_align_block_size' (#27311)

Signed-off-by: gnovack <gnovack@amazon.com>
Co-authored-by: Jee Jee Li <pandaleefree@gmail.com>
G
gnovack committed
8e4ca4d14e34305c0ab5640a9143871b25c91811
Parent: 1a0f4de
Committed by GitHub <noreply@github.com> on 10/22/2025, 12:23:57 PM