SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[CORE] Support Prefix Caching with Prompt Embeds (#27219)

Signed-off-by: Andrew Sansom <andrew@protopia.ai>
A
Andrew Sansom committed
ff93cc8c84b78effe5d096eea051e76e13927c5e
Parent: 243ed7d
Committed by GitHub <noreply@github.com> on 10/23/2025, 5:18:07 AM