COMMITS
May 14, 2026
A
fix: Modify release deletion command in workflow (#3307)
Alex Yang committed
J
[chore] add mamba codeowners list (#3318)
Jimmy Zhou committed
May 13, 2026
L
Fix: remove nvfp4 llama4 blocker (#3313)
Lain committed
S
Q
N
add cuda tile dependency for cuda 13.0 (#3305)
nv-yunzheq committed
I
Improved `simple` mamba SSU kernel (#2962)
Igor Shovkun committed
May 12, 2026
L
bench(moe_deepseek): scope autotune(True) to pre-warm only (#3301)
Lee Nau committed
K
J
ci: isolate nightly package tests from source tree (#3274)
Jonathan Dierksen committed
J
fix: remove over-strict K%4 assert in get_shuffle_matrix_sf_a_row_indices (#3163)
Jimmy Zhou committed
J
[chore] Add guard to blackwell GDN prefill (#3267)
Jiahan Chang (Cyrus) committed
May 11, 2026
V
fix(gdn_decode): widen pool indices to Int64 to prevent int32 element-offset overflow (#3230)
Vadim Gimpelson committed
L
bench(moe_deepseek): fix moe benchmark (supersedes #2886) (#3292)
Lee Nau committed
B
feat: add SM120 fmha_v2 kernels to AOT pip wheel builds (#2885)
Blake Ledden committed
A
fix(jit): propagate -DNDEBUG to host-side cflags (#3278)
Artem Perevedentsev committed
C
feat: FP8 output support for CUTLASS MLA paged attention (#2779)
Carl Y committed
M
Support Kimi K2.5 H64 CuTe DSL MLA decode (#3235)
Mingyang Wang committed
L
P
Add dynamic tokens-per-page TRTLLM-GEN GQA kernels (#3259)
Perkz Zheng committed
L
feat(moe): add SM120 W4A16 b12x kernels (#3271)
Luke Alonso committed
May 9, 2026
J
ci(jit-cache): limit sm110 builds to aarch64 (#3275)
Jonathan Dierksen committed
May 8, 2026
Y
non-override tactic control (#3260)
yanqinz2 committed
J
build: add sccache-backed jit-cache builds and AOT diagnostics (#3205)
Jonathan Dierksen committed
L
perf: optimize per-token nvfp4 quantization kernel. (#3237)
Lain committed
L
perf: Update moe gemm (#3239)
Lain committed
May 7, 2026
N
Loosened trtllm_ragged_attention_deepseek shape assertion (#3064)
nvjullin committed
K
Cutlass dsl 4.5 bump (#3246)
Ka-Hyun Nam committed
M
Issue #3047: Handle empty KV in MLA chunked-prefill (#3251)
Mingyang Wang committed
M
fix(sm12x): fix micro-kernel workspace sizing when routed_rows > num_local_experts (#3191)
meena-at-work committed