COMMITS
May 25, 2026
J
ggml : Parallelize quant LUT init (llama/23595)
Jeff Bolz committed
May 24, 2026
J
TP: fix entirely zero-sized slices per device (llama/23525)
Johannes Gäßler committed
S
opencl: batch profiling to improve speed and prevent memory leaks (llama/23495)
shaofeiqi committed
Y
hexagon: apply repl optimization in flash attn softmax as #22993 (llama/23455)
Yiwei Shao committed
May 23, 2026
D
J
vulkan: fix windows find_package of SPIRV-Headers (llama/23215)
Jeff Bolz committed
S
opencl: generalize Adreno MoE kernels on M (llama/23449)
Shawn Gu committed
May 22, 2026
A
SYCL: improve MoE prefill throughput (llama/23142)
Alexey Kopytko committed
A
sycl : Level Zero detection in ggml_sycl_init (llama/23097)
Alexey Kopytko committed
K
SYCL : gated_delta_net K>1 (llama/23174)
karavayev committed
K
SYCL: add BF16 to DMMV kernel path (~4x tg speedup on Intel Arc) (llama/21580)
Katostrofik committed
S
ggml-zendnn : add Q8_0 quantization support (llama/23414)
Sachin Sharma committed
May 21, 2026
J
CUDA: fix PDL CC check for JIT compilation (llama/23471)
Johannes Gäßler committed
P
C
fix(flash-attn): replace f32 with kv_type and q_type (llama/23372)
Chen Yuan committed
G
metal : optimize concat kernel and fix set kernel threads (llama/23411)
Georgi Gerganov committed
M
ggml : Check the right iface method before using the fallback 2d get (llama/23306)
Matt Corallo committed
T
hexagon: ssm-conv fix for large prompts (llama/23307)
Todor Boinovski committed
May 20, 2026
L
opencl: refactor backend initilization (llama/23318)
lhez committed
D
vulkan: optimize operations in the IM2COL shader (llama/22685)
Daniele committed
M
hexagon: HMX quantized matmul rework (llama/23368)
Max Krasnyansky committed
A
Programmatic Dependent Launch (PDL) for more performance on newer NVIDIA GPUs (Hopper+) (llama/22522)
Andreas Kieslinger committed
G
metal : optimize pad + cpy (llama/23354)
Georgi Gerganov committed
R
ggml-cuda: tune RDNA3 Q6_K MMVQ nwarps (llama/23349)
ravel7524 committed
May 19, 2026
S
opencl: add MoE support for q4_k, q5_k, q6_k on Adreno (llama/23303)
shaofeiqi committed
A
hexagon: add MROPE and IMROPE support in HTP rope op (llama/23317)
Aparna M P committed
A
hexagon: enable support for NORM op (llama/23319)
Aparna M P committed
R
ggml-webgpu : extend GDN for K>1 (llama/23299)
Reese Levine committed
I
sycl: add GGML_SYCL_USE_ASYNC_MEM_OP env toggle (llama/22153)
Intel AI Get-to Market Customer Success and Solutions committed
R
rpc : keep last_graph_uid in the device context (llama/23273)
Radoslav Gerganov committed