COMMITS
October 23, 2025
J
[Chore] Remove duplicate `has_` functions in vllm.utils (#27372)
Jonathan Chen committed
W
[Model] Add num_cached_tokens for PoolingRequestOutput (#27378)
wang.yuqi committed
G
[V1][spec decode] return logprobs for spec decoding (#26060)
Giancarlo Delfin committed
A
[CORE] Support Prefix Caching with Prompt Embeds (#27219)
Andrew Sansom committed
P
[Bugfix][Core] running queue index leakage exception (#26754)
PiteXChen committed
F
[Bugfix] Fix incorrect kv cache metrics in grafana.json (#27133)
fangpings committed
C
[Bugfix] Fix SLA tuner initialization (#27355)
Cyrus Leung committed
October 22, 2025
M
[MLA] Bump FlashMLA (#27354)
Matthew Bonanni committed
D
[Chore] Separate out system utilities from vllm.utils (#27201)
dongbo910220 committed
D
[BugFix] bugfix for Flash Attention MLA with full cuda graph IMA following pr-25490 (#27128)
Daisy-Ma-coder committed
R
[Feature] publisher default set zmq in kv_event config (#26915)
rongfu.leng committed
W
[Doc] Fix numbering sequence in prefix caching (#27357)
William Song committed
L
[Model] Revert PR #26715: Restore custom PaliGemma and Gemma3-MM impl… (#27309)
Luciano Martins committed
I
R
Support Anthropic API /v1/messages Endpoint (#22627)
RED committed
N
[P/D] Dynamic `kv_output_aggregator` collect size (#26734)
Nicolò Lucchesi committed
R
[Frontend] Require flag for loading text and image embeds (#27204)
Russell Bryant committed
I
[Bugfix] Fix HF format InternVL large variants video processing (#27330)
Isotr0py committed
C
[Bugfix] Make `get_mrope_input_positions` instance methods (#27342)
Cyrus Leung committed
C
[NIXL] use Host buffer to support TP_ratio > 1 for XPU (#27140)
Chendi.Xue committed
J
[Bugfix] Add missing 'is_internal_router' attribute to FusedMoEWithLoRA (#27351)
Jee Jee Li committed
R
[bugfix] remove unused parameters to reduce unnecessary vram usage (#26789)
Reinforce-II committed
W
[Bug] Fix DeepSeek-V2.5-1210-FP8 issue (#27267)
Wentao Ye committed
M
[NIXL] Terminate handshake listener thread in shutdown (#26404)
Mark McLoughlin committed
I
[Model] Upstream Deepseek-OCR model (#27247)
Isotr0py committed
D
[Chore] Separate out optional dependency checks from vllm.utils (#27207)
dongbo910220 committed
A
Mirroring changes in test-pipeline.yaml into test-amd.yaml (#27242)
Alexei-V-Ivanov-AMD committed
M
[docs] Update v1 metrics design doc (#27332)
Mark McLoughlin committed