SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Kernel] Add GPTQv2 format support for low-bit or asymmetric quantization, by adapting gptq_gemm (#26092)

X
Xiangyu Li committed
5cc6bddb6ef5e8e5c10de8122a43fd6e8c1e3b4b
Parent: 1f9460c
Committed by GitHub <noreply@github.com> on 10/24/2025, 3:26:13 AM