SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Bugfix] skip cuda graph for drafter when running with eager (#26821)

Signed-off-by: Benjamin Chislett <bchislett@nvidia.com>
B
Benjamin Chislett committed
19748806f04e6e390b42c2318a5ca76ffb6c1368
Parent: 4a8a567
Committed by GitHub <noreply@github.com> on 10/21/2025, 10:39:09 PM