feat(glm5-next): add GLM-5.3-Flash training support (#3699)
* feat(glm5-next): support GLM-5.3-Flash training Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * feat(glm5-next): add CP2 MedPix recipe Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * fix(glm5-next): keep CP gather live for empty shards Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * fix(glm5-next): keep validated EP72 CP2 recipe Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * fix: align GLM-5.3 checkpoint and MoE parity Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * chore(glm5-next): drop single-device loader change Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * fix(glm5-next): keep example W&B opt-in Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * perf(glm5-next): use cuDNN sparse attention Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * fix(glm5-next): guard optional imports and document model Keep the base installation importable without VLM or FLA extras, add regression coverage for the torchvision-free path, and register the new architecture in model coverage docs. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * fix(glm5-next): decay KDA state along key dimension Match the pure-Torch recurrent KDA fallback to FLA's reference state update and cover the decay axis numerically. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: HuiyingLi <willwin.lee@gmail.com> * fix(docs): place GLM-5.3 under publishing org Follow the existing zai-org to thudm documentation slug mapping so recipe documentation coverage recognizes the model card. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Signed-off-by: HuiyingLi <willwin.lee@gmail.com> --------- Signed-off-by: HuiyingLi <willwin.lee@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
H
Huiying committed
9228f33cf73d66a9b2e84256d298aac9a70283f0
Parent: 995898d
Committed by GitHub <noreply@github.com>
on 8/28/2026, 7:27:29 AM