SIGN IN SIGN UP

Chunked Prefill VLM (#3188)

* add logic

* working

* add encoder cache free

* fixes

* fix idefics

* update pixel_values

* add improvements

* add improvements

* improve

* nit

* fix inputs_embeds

* nit

* optimizations

* add prometheus port

* rename vars

* rename vars

* nit

* disable chunking for qwen

* review comments

* remove port

* improve headdim

* remove kwargs and redundant args

* fix qwen2_5

* fix config image_token_id error

* fix test

* update paligemma

* fix paligemma text

* minor fix

* fix qwen test

* fix qwen test
M
Mohit Sharma committed
329f612e55829adca0d36e88296e8b48990960b9
Parent: 533eee5
Committed by GitHub <noreply@github.com> on 5/6/2025, 4:01:59 PM