SIGN IN SIGN UP

A high-throughput and memory-efficient inference and serving engine for LLMs

0 0 122 Python

Granite 4.0 quark quantization support (#26944)

Signed-off-by: Xiao YU <Xiao.YU@xilinx.com>
Signed-off-by: Xiao Yu <xiao.yu.dc@outlook.com>
Co-authored-by: Xiao YU <Xiao.YU@xilinx.com>
X
xiao-llm committed
70022ffc002dabc59b936d3fc001b94b81ba08db
Parent: f417746
Committed by GitHub <noreply@github.com> on 10/24/2025, 2:14:03 AM