fix(distributed): avoid duplicate FSDP2 prefetch all-gathers (#3411)
* fix(distributed): use non-reentrant HF checkpointing Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(distributed): run FSDP2 prefetch regression in CI Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(distributed): cover HF checkpointing in robustness CI Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(distributed): preserve Llama robustness coverage Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): calibrate Llama cross-TP KL tolerance Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> --------- Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>
Y
Yuhe Zhang committed
99dc66552e63e36d5f8e06dcc89d7303ba918e2b
Parent: 2261cab
Committed by GitHub <noreply@github.com>
on 8/6/2026, 11:46:06 PM