COMMITS
/ examples/batched-bench/batched-bench.cpp October 18, 2024
X
llama : remove all_pos_0, all_pos_1, all_seq_id from llama_batch (#9745)
Xuan Son Nguyen committed
October 10, 2024
D
common : use common_ prefix for common library functions (#9805)
Diego Devesa committed
September 15, 2024
G
common : reimplement logging (#9418)
Georgi Gerganov committed
September 13, 2024
G
llama : llama_perf + option to disable timings during decode (#9355)
Georgi Gerganov committed
September 11, 2024
G
batched-bench : remove unused code (#9305)
Georgi Gerganov committed
September 9, 2024
X
common : move arg parser code to `arg.cpp` (#9388)
Xuan Son Nguyen committed
September 7, 2024
X
common : refactor arg parser (#9308)
Xuan Son Nguyen committed
G
llama : refactor sampling v2 (#9294)
Georgi Gerganov committed
September 6, 2024
A
batched-bench : add `--output-format jsonl` option (#9293)
Aarni Koskela committed
August 4, 2024
B
batched-bench : handle empty `-npl` (#8839)
Brian Cunnie committed
June 4, 2024
G
common : refactor cli arg parsing (#7675)
Georgi Gerganov committed
April 30, 2024
G
ggml : add Flash Attention (#5021)
Georgi Gerganov committed
April 5, 2024
T
bench : make n_batch and n_ubatch configurable in Batched bench (#6500)
Ting Sun committed
March 13, 2024
S
llama : add pipeline parallelism support (#6017)
slaren committed
March 11, 2024
G
llama : more consistent names of count variables (#5994)
Georgi Gerganov committed
March 8, 2024
C
llama : support Mamba Selective State Space Models (#5328)
compilade committed
March 1, 2024
P
llama : cleanup unused mmq flags (#5772)
Pierrick Hymbert committed
February 18, 2024
H
ggml, common, examples, tests : fixed type arguments in printf (#5528)
Herman Semenov committed
February 16, 2024
B
ggml : add numa options (#5377)
bmwl committed
January 31, 2024
G
llama : remove LLAMA_MAX_DEVICES and LLAMA_SUPPORTS_GPU_OFFLOAD (#5240)
Georgi Gerganov committed
January 12, 2024
S
llama : ggml-backend integration (#4766)
slaren committed
December 1, 2023
G
ggml : add ggml_soft_max_ext (#4256)
Georgi Gerganov committed
October 29, 2023
K
Extend llama_kv_cache_seq_rm to allow matching any sequence (#3843)
Kerfuffle committed
October 25, 2023
G
batched-bench : print params at start
Georgi Gerganov committed
October 18, 2023
G
speculative : add tree-based sampling example (#3624)
Georgi Gerganov committed
October 11, 2023
G
batched : add bench tool (#3545)
Georgi Gerganov committed