SIGN IN SIGN UP

test(checkpoint): expand parity metrics and phase coverage (#3567)

* test(checkpoint): expand parity metrics and phase coverage

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): calibrate all parity phases on long input

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): use unique long-context parity input

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): stabilize long-context parity input

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): define representative parity cohort

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): calibrate parity profiles

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): relax Step resume drift

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): calibrate parity profiles from scoped CI

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): capture checkpoint repeatability metrics

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* docs(ci): explain repeatability metric diagnostics

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): add high-variance checkpoint parity profile

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): avoid redundant resume checkpoint writes

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* refactor(checkpoint): use targeted Step parity overrides

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* fix(test): repair checkpoint parity CI regressions

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(checkpoint): cover Nemotron Flash HF reload

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* fix(test): stabilize remote checkpoint parity

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* fix(ci): use cached Nemotron family tokenizer

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): enable routed MoE resume coverage

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* fix(checkpoint): map Nemotron PEFT export namespace

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* fix(ci): reach configured resume boundary

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): calibrate Nemotron chat resume drift

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* docs(ci): clarify checkpoint robustness compatibility

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): enable Nemotron resume coverage

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): preserve Nemotron tokenizer issue gate

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): retire stale Nemotron KL overrides

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): forbid legacy max KL recipe thresholds

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): remove unmeasured Mistral FP8 overrides

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): remove legacy checkpoint parity fields

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): unify checkpoint parity overrides

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* test(ci): support per-comparison parity profiles

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

* fix(ci): keep parity threshold numeric across YAML loaders

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>

---------

Signed-off-by: Yuhe Zhang <yuhez@nvidia.com>
Co-authored-by: Alexandros Koumparoulis <153118171+akoumpa@users.noreply.github.com>
Y
Yuhe Zhang committed
63955a912fe9d1cc1e498c9ea2e387b6c366fca1
Parent: ad1c7d4
Committed by GitHub <noreply@github.com> on 8/25/2026, 8:34:20 PM