COMMITS
August 30, 2026
A
A
hybrid: prefill cached follow-ups above 4*decode
andi committed
August 27, 2026
A
server: skip wide prefill when prompt cache will continue
andi committed
A
hybrid: add warm replace-ratio hysteresis
andi committed
August 25, 2026
A
server: continue GDN prompt cache across retokenized history
andi committed
A
chore: add LFM2.5 and Ornith GGUF download scripts
andi committed
A
August 24, 2026
A
hybrid: reuse expert prefetch staging across prompts
andi committed
August 23, 2026
A
hybrid: keep DFlash KV, GDN reuse, and on-demand router
andi committed
August 20, 2026
A
A
cuda: port fixes and add Gemma 4 launcher
andi committed
August 19, 2026
A
fix: resolve dependency security alerts
andi committed
A
perf: optimize Ling-tiny CUDA routing
andi committed
A
feat: add Ling-tiny starter for GTX 1080
andi committed
August 18, 2026
A
feat: enable Ling phase-batching without cpu-moe
andi committed
A
cpu: use ggml_cpu_fp16_to_fp32 in tiled flash attention
andi committed
A
revert: drop OpenWebUI-specific think-block stripping
andi committed
A
feat: add TurboLLM profiles for GTX 1660 Ti and GTX 1080
andi committed
A
feat: add bailingmoe3 support for Ling-3.0
andi committed
A
feat: allow KVFlash paging on MLA K-only caches
andi committed
August 17, 2026
A
perf: shrink prefill graphs and retune start1660 for 1660 Ti
andi committed
August 14, 2026
A
hybrid : keep DFlash draft prefix after GDN rebuild
andi committed
A
hybrid : fix turn-2 reasoning budget and LCP prompt reuse
andi committed
August 13, 2026
A
A
Merge KVFlash paging overlap and QK residency
andi committed
A
perf: overlap KVFlash paging and add target-QK residency
andi committed
August 12, 2026
A
Merge KVFlash prompt-cache state and reuse
andi committed
A
feat: enable KVFlash prompt cache state and reuse
andi committed
R
Merge origin/main with DDTree state commit
root committed
A
Merge KVFlash paging hardening
andi committed