COMMITS
August 21, 2026
I
sycl: fix multiple warnings in compiling sycl backend (#26713)
Ian Faust committed
N
sycl : fix load model with mlock issue (#27250)
Neo Zhang committed
C
TP: enable tensor split for LFM2/LFM2MOE (#26993)
Chris Danis committed
V
docs: fix typos in ET.md (#27457)
vk committed
August 20, 2026
X
ggml: support ggml_rope_set_offset on opencl, sycl, wgpu, hexagon (#27345)
Xuan-Son Nguyen committed
E
ci: use shell script to check cmake pkg (#27414)
Eve committed
G
metal : clamp K extent in tensor API mat-mat kernel for K not a multiple of 32 (#27450)
Georgi Gerganov committed
H
opencl: fix q6_K flat mul_mat for Adreno A6x/A7x GPUs with older E031 compilers (#26476)
Hongqiang Wang committed
L
opencl: fix local size for norm (#27339)
lhez committed
A
ui: Stores split refactor (#27240)
Aleksander Grygier committed
J
mtmd: add --mmproj-device argument (#23255)
John-Henry Lim committed
T
model : support DSpark for LFM2 models (#27383)
Tarek Dakhran committed
J
vulkan: FA MMQ should use fp32 for Q quantization calculations (#27413)
Jeff Bolz committed
G
metal : dequant kv cache only for large batches (#27438)
Georgi Gerganov committed
O
CI: Use LLVM's OpenMP over MSVC_DEBUG_non_redist on Windows (#26678)
Oliver Simons committed
X
server: (router) lazy-load startup_models after main setup (#27424)
Xuan-Son Nguyen committed
A
server : fix --docker-repo being treated as router mode (#27416)
Aritro Bandyopadhyay committed
P
CUDA: adding switch points per HW and quant type to tune the mvq->MMQ decode crossover (#26079)
Pranesh Gonegandla committed
A
common : gracefully fallback on unsupported regex patterns in JSON schema (#26939)
Aldehir Rojas committed
G
metal : dequantize quantized KV to F16 before flash attention (#27390)
Georgi Gerganov committed
G
Revert "tensor-split meta backend fixes (#26502)" (#27433)
Georgi Gerganov committed
R
ggml: fix backend split scheduler race condition (#26040)
Ruben Ortlam committed
R
convert: fix get block count error for Nemotron 3 Ultra (#27101)
Rock Chen committed
A
ggml-cuda: provide static workspace for cuBLAS handles (#26574)
Alexander Heisler committed
G
graph : create V as a view of K in the k_iswa build_attn (#27392)
Georgi Gerganov committed
G
spec : avoid binding reference to null pointer (#27404)
Georgi Gerganov committed
M
vulkan : add source groups for shaders (#26666)
Markus Tavenrath committed
H
opencl: make the MoE expert scatter deterministic (#26464)
Hongqiang Wang committed
August 19, 2026
M
tensor-split meta backend fixes (#26502)
Max Krasnyansky committed
Y
hexagon: fix FA HMX queue ordering and pack the rescale D matrices (#27042)
Yiwei Shao committed