SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

vllm bench serve shows num of failed requests (#26478)

Signed-off-by: Tomas Ruiz <tomas.ruiz.te@gmail.com>
T
Tomas Ruiz committed
965c5f4914595a96d3c2d6ee2d0ed9dd00a8d573
Parent: 4d055ef
Committed by GitHub <noreply@github.com> on 10/17/2025, 2:55:09 AM