SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Distributed] Basic set of configuration for large EP deployment on GB200 (#27328)

Signed-off-by: Pengchao Wang <wpc@fb.com>
Co-authored-by: Lu Fang <30275821+houseroad@users.noreply.github.com>
P
Pengchao Wang committed
d95d0f4b985f28ea381e301490f9d479b34d8980
Parent: 0402428
Committed by GitHub <noreply@github.com> on 10/24/2025, 9:16:44 PM