COMMITS
August 31, 2026
A
fix(generate): chunk prefill for prompts shorter than prefill_step_size (#2119)
Alazer Manakelew committed
A
fix(speculative): detect native MTP drafters via num_nextn_predict_layers (#2120)
Alazer Manakelew committed
A
muse_glimmer: load flattened OptiQ checkpoints (#2108) (#2118)
Alazer Manakelew committed
A
Add LongCat-Flash-Lite-Sparse (LSA + n-gram) (#2063)
Alazer Manakelew committed
A
Add DeepSeek-V4 DSpark speculative drafter (fixes #2095) (#2097)
Alazer Manakelew committed
A
Clamp ragged batched speculative acceptance to fix phantom KV (#1962) (#2113)
Alazer Manakelew committed
A
Add LLaVA-OneVision (#2115)
Alazer Manakelew committed
A
fix: stop generation on the tokenizer's EOS when the config omits it (#2112)
Alazer Manakelew committed
L
Merge pull request #2110 from Thump604/604/nemotron-audio-dependency-floor
Lucas Newman committed
August 30, 2026
T
fix: require compatible mlx-audio for VoiceChat
Thump604 committed
A
diffusion_gemma: strip channel scaffolding from decoded output (#2101)
Alazer Manakelew committed
A
fix: load float-quantized compressed-tensors checkpoints (#2102)
Alazer Manakelew committed
A
fix(video): resolve frame sampling in one place (#2050) (#2100)
Alazer Manakelew committed
August 29, 2026
A
lfm2_vl: load flattened OptiQ checkpoints (#2088) (#2092)
Alazer Manakelew committed
August 28, 2026
Y
Fix TurboQuant batch decode and speculative KV
Yuhua Chen committed
A
Fix server model discovery scope (#2076)
Alazer Manakelew committed
L
Merge pull request #2045 from Thump604/604/qwen38-external-ple-upstream
Lucas Newman committed
L
Merge pull request #2046 from Thump604/604/qwen38-adaptive-mtp
Lucas Newman committed
L
Merge branch 'main' into 604/qwen38-external-ple-upstream
Lucas Newman committed
L
fix: preserve partial quantization overrides
Lucas Newman committed
P
Bump version to 0.7.0rc0 (#2083)
Prince Canuma committed
L
Merge pull request #2075 from Anai-Guo/fix-recurrent-gemma-module-init
Lucas Newman committed
L
Merge pull request #2036 from Anai-Guo/fix-chunked-kv-trim-valid-length
Lucas Newman committed
A
moe-offloading (#1813)
Alazer Manakelew committed
P
Redesign APC for dense and hybrid model caches (#1960)
Prince Canuma committed
T
Fix Qwen4 MTP layer type metadata
Thump604 committed
A
Add pr contributor reminder (#2068)
Alazer Manakelew committed
A
Merge branch 'main' into 604/qwen38-external-ple-upstream
Alazer Manakelew committed
A
Fix nemotron_h prefill crash when inputs and inputs_embeds are both passed (#2057)
Alazer Manakelew committed