SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Misc] Add TPU usage report when using tpu_inference. (#27423)

Signed-off-by: Hongmin Fan <fanhongmin@google.com>
H
hfan committed
8dbe0c527fa76cd908bb6287f9f501df44e04473
Parent: 5cc6bdd
Committed by GitHub <noreply@github.com> on 10/24/2025, 3:29:37 AM