SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

[Feature][Quantization] auto_round support for mixed bits quantization (#23812)

Signed-off-by: n1ck-guo <heng.guo@intel.com>
Signed-off-by: Heng Guo <heng.guo@intel.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
H
Heng Guo committed
87778d5f00ab09f88d0b677c34ae7d33e6296ce2
Parent: f9e7ad5
Committed by GitHub <noreply@github.com> on 10/20/2025, 10:23:30 PM