SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[CI/Build] tests(v1): feed Triton attention the (num_blocks, 2, …) KV cache layout in backend-correctness tests (#26663)

Signed-off-by: Huamin Li <3ericli@gmail.com>
Co-authored-by: Ye (Charlotte) Qi <yeq@meta.com>
H
Huamin Li committed
c312320764193e7d0ffa99d247c61efe5458a635
Parent: c981f0e
Committed by GitHub <noreply@github.com> on 10/18/2025, 4:11:26 AM