COMMITS
/ examples/quantize/quantize.cpp December 7, 2024
D
ggml : refactor online repacking (#10446)
Djip007 committed
September 20, 2024
S
quantize : improve type name parsing (#9570)
slaren committed
September 6, 2024
C
ggml-quants : ternary packing for TriLMs and BitNet b1.58 (#8151)
compilade committed
August 24, 2024
J
quantize : fix typo in usage help of `quantize.cpp` (#9145)
João Dinis Ferreira committed
August 6, 2024
D
quantize : update usage comment in quantize.cpp (#8889)
Daniel Bevenius committed
July 16, 2024
G
llama : valign + remove unused ftype (#8502)
Georgi Gerganov committed
July 10, 2024
D
ggml : add AArch64 optimized GEMV and GEMM Q4 kernels (#5780)
Dibakar Gope committed
June 22, 2024
May 22, 2024
G
common : normalize naming style (#7462)
Georgi Gerganov committed
May 19, 2024
F
quantize : fix --keep-split check (#7374)
Fred Douglas committed
May 8, 2024
J
ggml : introduce bfloat16 support (#6412)
Justine Tunney committed
April 26, 2024
P
quantize: add imatrix and dataset metadata in GGUF (#6658)
Pierrick Hymbert committed
April 25, 2024
April 3, 2024
S
ggml : mul_mat_id use the same tensor for all the experts (#6387)
slaren committed
March 26, 2024
K
IQ1_M: 1.75 bpw quantization (#6302)
Kawrakow committed
K
quantize : be able to override metadata by key (#6321)
Kawrakow committed
March 22, 2024
K
quantize: options for output and token embedding tensors qtype (#6239)
Kawrakow committed
February 27, 2024
K
IQ4_XS: a 4.25 bpw quantization (#5747)
Kawrakow committed
February 26, 2024
February 24, 2024
K
IQ3_S: a much better alternative to Q3_K (#5676)
Kawrakow committed
February 21, 2024
K
IQ4_NL: 4-bit non-linear quants with blocks of 32 (#5590)
Kawrakow committed
February 18, 2024
K
1.5 bit quantization (#5453)
Kawrakow committed
February 16, 2024
B
ggml : add numa options (#5377)
bmwl committed
February 3, 2024
M
refactor : switch to emplace_back to avoid extra object (#5291)
Michael Klimenko committed
January 30, 2024
K
SOTA 3-bit quants (#5196)
Kawrakow committed
V
quantize : fix typo (#5211)
Vladimir Malyutin committed
January 22, 2024
K
llama : add Q3_K_XS (#5060)
Kawrakow committed
January 14, 2024
K
Add ability to use importance matrix for all k-quants (#4930)
Kawrakow committed
K
2-bit quantizations (#4897)
Kawrakow committed
January 11, 2024
K
llama : restore intended k-quants mixes for MoE models (#4872)
Kawrakow committed
November 2, 2023
C
build : link against build info instead of compiling against it (#3879)
cebtenzzre committed
October 29, 2023
G
ggml : quantization refactoring (#3833)
Georgi Gerganov committed
September 28, 2023
C
build : enable more non-default compiler warnings (#3200)
Cebtenzzre committed
September 18, 2023
C
make : restore build-info.h dependency for several targets (#3205)
Cebtenzzre committed
September 15, 2023
C
examples : add compiler version and target to build info (#2998)
Cebtenzzre committed
C
check C++ code with -Wmissing-declarations (#3184)
Cebtenzzre committed
September 7, 2023
C
fix some warnings from gcc and clang-tidy (#3038)
Cebtenzzre committed
September 1, 2023
K
Allow quantize to only copy tensors, some other improvements (#2931)
Kerfuffle committed
August 28, 2023
C
quantize : make output filename optional again (#2823)
Cebtenzzre committed
August 23, 2023
K
Fix values shown in the quantize tool help (#2735)
Kawrakow committed
August 21, 2023
G
gguf : new file format with flexible meta data (beta) (#2398)
Georgi Gerganov committed
July 18, 2023
G
llama : shorten quantization descriptions
Georgi Gerganov committed
July 10, 2023
E
mpi : add support for distributed inference via MPI (#2099)
Evan Miller committed
June 26, 2023
Z
ggml : add NUMA support (#1556)
zrm committed
June 13, 2023
K
Allow "quantizing" to f16 and f32 (#1787)
Kerfuffle committed
June 10, 2023
June 5, 2023
K
ggml : add SOTA 2,3,4,5,6 bit k-quantizations (#1684)
Kawrakow committed
May 20, 2023
G
llama : add llama_init_backend() API (close #1527)
Georgi Gerganov committed
May 11, 2023
G
ggml : remove bit shuffling (#1405)
Georgi Gerganov committed
May 4, 2023