SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Core][Hybrid allocator + kv connector 1/n] Enable hybrid allocator + KV cache connector (#25712)

Signed-off-by: KuntaiDu <kuntai@uchicago.edu>
Signed-off-by: Kuntai Du <kuntai@uchicago.edu>
K
Kuntai Du committed
b853540388bf8c796e1c88dabceaeebaab73621e
Parent: 56ed760
Committed by GitHub <noreply@github.com> on 10/25/2025, 6:34:18 AM