COMMITS
/ examples/llama-bench/llama-bench.cpp November 25, 2024
D
ggml : add support for dynamic loading of backends (#10469)
Diego Devesa committed
November 20, 2024
D
llama : add .clang-format file (#10415)
Diego Devesa committed
November 14, 2024
D
ggml : build backends as libraries (#10256)
Diego Devesa committed
November 8, 2024
G
metal : optimize FA kernels (#10171)
Georgi Gerganov committed
October 30, 2024
D
llama : refactor model loader with backend registry (#10026)
Diego Devesa committed
October 18, 2024
X
llama : remove all_pos_0, all_pos_1, all_seq_id from llama_batch (#9745)
Xuan Son Nguyen committed
O
[SYCL] Add SYCL Backend registry, device and Event Interfaces (#9705)
Ouadie EL FAROUKI committed
October 10, 2024
D
rpc : add backend registry / device interfaces (#9812)
Diego Devesa committed
September 17, 2024
M
llama-bench: correct argument parsing error message (#9524)
Michael Podvitskiy committed
September 13, 2024
G
llama : llama_perf + option to disable timings during decode (#9355)
Georgi Gerganov committed
September 7, 2024
G
llama : refactor sampling v2 (#9294)
Georgi Gerganov committed
September 6, 2024
A
llama-bench : log benchmark progress (#9287)
Aarni Koskela committed
September 5, 2024
S
llama-bench : fix NUL terminators in CPU name (#9313)
slaren committed
September 4, 2024
R
rpc : make RPC servers come first in the device list (#9296)
Radoslav Gerganov committed
September 3, 2024
A
llama-bench : add JSONL (NDJSON) output mode (#9288)
Aarni Koskela committed
August 29, 2024
F
Threadpool: take 2 (#8672)
Faisal Zaghloul committed
August 7, 2024
Z
llama-bench : add support for getting cpu info on Windows (#8824)
Zhenwei Jin committed
July 27, 2024
S
ggml : reduce hash table reset cost (#8698)
slaren committed
July 17, 2024
H
[CANN] Add Ascend NPU backend (#6035)
hipudding committed
June 14, 2024
R
llama-bench : fix RPC indication (#7936)
Radoslav Gerganov committed
June 13, 2024
S
move BLAS to a separate backend (#6210)
slaren committed
June 11, 2024
J
llama-bench: more compact markdown tables (#7879)
Johannes Gäßler committed
June 4, 2024
G
common : refactor cli arg parsing (#7675)
Georgi Gerganov committed
G
ggml : remove OpenCL (#7735)
Georgi Gerganov committed
S
May 29, 2024
R
llama-bench : add support for the RPC backend (#7435)
Radoslav Gerganov committed
May 22, 2024
G
common : normalize naming style (#7462)
Georgi Gerganov committed
S
phi3 : duplicate rope factors in each layer (#7447)
slaren committed
May 10, 2024
S
llama-bench : add pp+tg test type (#7199)
slaren committed
May 5, 2024
K
Adding support for the --numa argument for llama-bench. (#7080)
kunnis committed
April 30, 2024
G
ggml : add Flash Attention (#5021)
Georgi Gerganov committed
April 16, 2024
J
ggml : add llamafile sgemm (#6414)
Justine Tunney committed
March 26, 2024
S
cuda : rename build flag to LLAMA_CUDA (#6299)
slaren committed
March 21, 2024
K
Add ability to use Q5_0, Q5_1, and IQ4_NL for quantized K cache (#6183)
Kawrakow committed
March 18, 2024
S
backend : offload large batches to GPU (#6083)
slaren committed
March 15, 2024
S
March 14, 2024
S
gguf : fix resource leaks (#6061)
Steve Grubb committed
March 13, 2024
S
llama : add pipeline parallelism support (#6017)
slaren committed
March 7, 2024
G
llama-bench : add embeddings option (#5924)
Georgi Gerganov committed
March 2, 2024
N
Support multiple GPUs (split mode) on SYCL backend (#5806)
Neo Zhang Jianyu committed
March 1, 2024
P
llama : cleanup unused mmq flags (#5772)
Pierrick Hymbert committed
February 25, 2024
G
code : normalize enum names (#5697)
Georgi Gerganov committed
February 16, 2024
B
ggml : add numa options (#5377)
bmwl committed
February 3, 2024
M
refactor : switch to emplace_back to avoid extra object (#5291)
Michael Klimenko committed
February 1, 2024
N
add --no-mmap in llama-bench (#5257)
Neo Zhang Jianyu committed
January 31, 2024
G
llama : remove LLAMA_MAX_DEVICES and LLAMA_SUPPORTS_GPU_OFFLOAD (#5240)
Georgi Gerganov committed
J
kompute : llama-bench support and ggml_cpu_has_kompute() (#5226)
Jared Van Bortel committed
January 28, 2024
0
ggml : add Vulkan backend (#2059)
0cc4m committed
January 12, 2024
S
llama : ggml-backend integration (#4766)
slaren committed
January 7, 2024
S
llama-bench : add no-kv-offload parameter (#4812)
slaren committed