perf(checkpoint): reduce allocating grouped MoE load overhead (#3580)
* perf(checkpoint): reduce grouped MoE fallback overhead Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * perf(checkpoint): avoid initialized load destination copies Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * perf(checkpoint): load Gemma4 experts into final storage Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * perf(checkpoint): load Gemma4 EP shards directly Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * ci(gemma4): reduce robustness time limit Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * fix(checkpoint): reuse contiguous TE load views Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * refactor(checkpoint): reuse load views without blank buffers Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * docs(checkpoint): clarify load tensor terminology Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * fix(lint): remove invalid checkpoint suppression Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> --------- Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> Co-authored-by: Alexandros Koumparoulis <153118171+akoumpa@users.noreply.github.com>
Y
Yuhe Zhang committed
ad1c7d4cfbed1a89ee5f392cbdc84860c09d4615
Parent: 1ae556e
Committed by GitHub <noreply@github.com>
on 8/25/2026, 5:01:52 PM