COMMITS
August 30, 2026
A
ci(build): bump FA4, quack and ffpa-attn (#3733)
Alexandros Koumparoulis committed
August 29, 2026
H
docs(glm): add GLM-5.3 training coverage (#3752)
Huiying committed
H
feat(glm): add GLM-5.3 cuDNN training support (#3747)
Huiying committed
Y
fix(moe): Kimi-K3 correctness fixes (routing, NCCL warmup, PP, per-GPU tps) (#3731)
yisongbetter committed
H
docs(models): expand GLM-5.3-Flash coverage (#3744)
Huiying committed
August 28, 2026
Y
perf(qwen3.5): reuse packed GDN metadata across layers (#3705)
Yuhe Zhang committed
Y
fix(glm): align and diagnose cross-framework router parity (#3635)
Yuhe Zhang committed
A
test(pp): add qwen3_5_moe PP and EP parity tests [3/4] (#3649)
Abhishree Thittenamane committed
P
fix(mtp): enable activation checkpointing for repeated dense MTP blocks (#3660)
Piotr Żelasko committed
H
feat(glm5-next): add GLM-5.3-Flash training support (#3699)
Huiying committed
A
fix: catch FileExistsError when checkpointing to distributed fs (#3686)
Aarni Koskela committed
S
ci: AUT-2161 route external jobs through ephemeral v2 (#3697)
svcnemo-autobot committed
August 27, 2026
S
fix(ci): AUT-2165 serialize approval queue runs (#3712)
svcnemo-autobot committed
S
ci: AUT-2165 trigger approval queue when CICD completes (#3710)
svcnemo-autobot committed
A
test(pp): fail parity tests when neither run trained (#3701)
Abhishree Thittenamane committed
Y
fix(minimax): repair HF parity reference and gate at the measured envelope (#3674)
Yuhe Zhang committed
H
feat(qwen3-8-flash-next): add Engram training (#3690)
Huiying committed
H
fix(fsdp): preserve internal module output dtype (#3657)
Huiying committed
A
ci: merge L2 matrix entries so the e2e stage fits max-parallel: 10 (#3695)
Alexandros Koumparoulis committed
August 26, 2026
A
ci(convergence): fix gemma4 gate tolerance (#3696)
Abhishree Thittenamane committed
O
fix(retrieval): register Llama classifier auto class (#2905)
Oliver Holworthy committed
H
docs(models): add Qwen3.8-Flash-Next coverage (#3691)
Huiying committed
A
fix(peft): support mixed-dtype memory-efficient LoRA backward (#3675)
Alexandros Koumparoulis committed
H
A
docs(examples): add Qwen3-32B SWE-bench Verified eval (Phase 3) (#3671)
Abhishree Thittenamane committed
Y
fix(test): unblock the Kimi-Linear vanilla-HF parity reference (#3659)
Yuhe Zhang committed
August 25, 2026
Y
test(ci): report Jensen-Shannon checkpoint parity metrics (#3620)
Yuhe Zhang committed
A
chore: bump lint tools + fix lint issues (#3634)
Aarni Koskela committed
Y
perf(checkpoint): avoid full GC scans during MoE export (#3621)
Yuhe Zhang committed
Y
test(checkpoint): expand parity metrics and phase coverage (#3567)
Yuhe Zhang committed