COMMITS
August 25, 2026
Y
perf(checkpoint): reduce allocating grouped MoE load overhead (#3580)
Yuhe Zhang committed
A
docs(examples): add Gemma4-31B SWE-bench eval steps + base-vs-SFT agentic result (#3585)
Abhishree Thittenamane committed
August 24, 2026
D
fix(deps): resolve 26.08 rc10 container CVEs (#3647)
Dong Hyuk Chang committed
K
feat(dllm): add SCDD training and sampling (#3354)
Kashif Rasul committed
A
test(pp): add deepseek_v4 PP and EP parity tests [2/4] (#3613)
Abhishree Thittenamane committed
S
refactor(checkpoint): StateDictAdapter method target-module rename (#3591)
Stanley Mei committed
A
fix(kimi_k25): map lm_head to model.language_model.lm_head (#3632)
Aarni Koskela committed
A
feat(laguna): add Laguna XS 2.1 training recipe (#3639)
Alexandros Koumparoulis committed
August 23, 2026
A
fix: EP collective deadlock with variable-length token counts (LoRA flavor) (#3631)
Aarni Koskela committed
K
K
feat(streaming): async producer + SharedDirFeatureStore (#3094)
Kashif Rasul committed
August 22, 2026
K
feat(speculative): add Kimi K3 DFlash draft training (#3287)
khazzz1c committed
R
fix(checkpoint): load base weights under DDP and MegatronFSDP (#3597)
Roman Ralovets committed
A
fix: remove unused RUN_ID env var from CICD summary job (#3536)
Andrew White committed
S
fix(ci): AUT-1630 route autobot through internal queue (#3629)
svcnemo-autobot committed
R
fix(perf): recognize RTX PRO 6000 for MFU (#3598)
Roman Ralovets committed
K
docs: announce DFlash 2 support in the README (#3624)
khazzz1c committed
K
feat(dflash): add DFlash 2 draft model, trainer, and recipe (#3605)
Kashif Rasul committed
August 21, 2026
A
fix(deps): align pyproject DeepEP pin with the container build (#3608)
Alexandros Koumparoulis committed
D
fix(deps): resolve 26.08 rc9 container CVEs (#3607)
Dong Hyuk Chang committed
A
fix(lint): update Qwen3 prefilter optional annotations (#3602)
Alexandros Koumparoulis committed
A
chore(lint): enforce PEP 604 optional annotations (#3535)
Alexandros Koumparoulis committed
August 20, 2026
H
feat(glm): add cuDNN DSA backend for GLM-5.2 (#3127)
Huiying committed
A
feat(examples): Qwen3-32B CoderForge long-context CP data + SFT recipe (#3594)
Abhishree Thittenamane committed
Z
fix(diffusion): apply compute_dtype via autocast when FSDP2 skips (#3589)
Zeyu Zhou committed
A
test(pp): add gemma4 parallelism parity tests [1/4] (#3529)
Abhishree Thittenamane committed
Y
feat(vlm): add CP vision shape floor (#3581)
Yuhe Zhang committed
N
perf(Gemma4): reduce comms in sliding window, FA implementation, (#2984)
Nathan Azrak committed
A
fix(vlm): truncate only the token axis in gemma4 collate (#3531)
Abhishree Thittenamane committed
A
fix(pp): keep accumulated grads across stage re-initialization (#3530)
Abhishree Thittenamane committed