SIGN IN SIGN UP

perf(checkpoint): reduce allocating grouped MoE load overhead (#3580)

* perf(checkpoint): reduce grouped MoE fallback overhead

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* perf(checkpoint): avoid initialized load destination copies

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* perf(checkpoint): load Gemma4 experts into final storage

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* perf(checkpoint): load Gemma4 EP shards directly

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* ci(gemma4): reduce robustness time limit

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* fix(checkpoint): reuse contiguous TE load views

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* refactor(checkpoint): reuse load views without blank buffers

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* docs(checkpoint): clarify load tensor terminology

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* fix(lint): remove invalid checkpoint suppression

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

---------

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>
Co-authored-by: Alexandros Koumparoulis <153118171+akoumpa@users.noreply.github.com>
Y
Yuhe Zhang committed
ad1c7d4cfbed1a89ee5f392cbdc84860c09d4615
Parent: 1ae556e
Committed by GitHub <noreply@github.com> on 8/25/2026, 5:01:52 PM