SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Attention] Add MLA prefill backend: trtllm_ragged_attention_deepseek (#26397)

Signed-off-by: Ming Yang <minos.future@gmail.com>
M
Ming Yang committed
0f67d4d962872767ac1fca8e98d1bb679aae762a
Parent: 7e1d697
Committed by GitHub <noreply@github.com> on 10/24/2025, 5:24:08 PM