test(checkpoint): expand parity metrics and phase coverage (#3567)
* test(checkpoint): expand parity metrics and phase coverage Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): calibrate all parity phases on long input Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): use unique long-context parity input Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): stabilize long-context parity input Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): define representative parity cohort Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): calibrate parity profiles Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): relax Step resume drift Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): calibrate parity profiles from scoped CI Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): capture checkpoint repeatability metrics Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * docs(ci): explain repeatability metric diagnostics Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): add high-variance checkpoint parity profile Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): avoid redundant resume checkpoint writes Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * refactor(checkpoint): use targeted Step parity overrides Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * fix(test): repair checkpoint parity CI regressions Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(checkpoint): cover Nemotron Flash HF reload Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * fix(test): stabilize remote checkpoint parity Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * fix(ci): use cached Nemotron family tokenizer Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): enable routed MoE resume coverage Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * fix(checkpoint): map Nemotron PEFT export namespace Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * fix(ci): reach configured resume boundary Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): calibrate Nemotron chat resume drift Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * docs(ci): clarify checkpoint robustness compatibility Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): enable Nemotron resume coverage Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): preserve Nemotron tokenizer issue gate Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): retire stale Nemotron KL overrides Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): forbid legacy max KL recipe thresholds Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): remove unmeasured Mistral FP8 overrides Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): remove legacy checkpoint parity fields Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): unify checkpoint parity overrides Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * test(ci): support per-comparison parity profiles Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> * fix(ci): keep parity threshold numeric across YAML loaders Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> --------- Signed-off-by: Yuhe Zhang <yuhez@nvidia.com> Co-authored-by: Alexandros Koumparoulis <153118171+akoumpa@users.noreply.github.com>
Y
Yuhe Zhang committed
63955a912fe9d1cc1e498c9ea2e387b6c366fca1
Parent: ad1c7d4
Committed by GitHub <noreply@github.com>
on 8/25/2026, 8:34:20 PM