gemma : use more bits for the token_embd.weight tensor (#5650)
* gemma : use Q8_0 for the token_embd.weight tensor * llama : quantize token_embd.weight using output type
G
Georgi Gerganov committed
96633eeca1265ed03e57230de54032041c58f9cd
Parent: 847eedb
Committed by GitHub <noreply@github.com>
on 2/22/2024, 9:23:46 PM