COMMITS
October 21, 2025
N
[Nixl] Minor refactor to handshake related metadata (#26410)
Nicolò Lucchesi committed
Z
[Fix][Spec Decode] Fix llama4 draft loading with different quantization (#27136)
Zebing Lin committed
B
[Bugfix] Fix broken MTP weight loading for FP8 KV Scales (#27227)
Benjamin Chislett committed
V
[Bugfix] Fix gpt-oss w4a8 DP/EP on B200 (#26729)
Varun Sundar Rabindranath committed
S
[ModelOpt] Load w13/w2_input_scale for all experts, nvfp4 (#26135)
Shu Wang committed
P
[BugFix] GPT-OSS Attention DP + MoE TP weight loading issue (#24032)
Po-Han Huang (NVIDIA) committed
C
[Feature][Kernel]FusedMoE LoRA (#21229)
Chen Wu committed
R
[Frontend] Enforce tokenize=False when applying chat template (#27205)
Russell Bryant committed
L
create is_in_the_same_node on cpu (#26832)
Lunwen He committed
F
[cpu] Dispatch un-quantized linear to oneDNN/ACL by default for AArch64 (#27183)
Fadi Arafeh committed
N
[V0 Deprecation] Remove V0 metrics code (#27215)
Nick Hill committed
I
A
[ez] add uv lock to gitignore (#27212)
Andrew Xia committed
C
[ROCm] enable some tests in entrypoints test groups on AMD (#26725)
Concurrensee committed
October 20, 2025
H
N
[Bugfix][CI] Fix `Distributed Tests (4 GPUs)` async_sched+ray test (#27195)
Nicolò Lucchesi committed
S
E
Nemotron Nano V2 VL + EVS Video Support (#27107)
Eugene Khvedchenya committed
I
AArch64 CPU Docker pipeline (#26931)
ioana ghiban committed
J
[Kernel] Accelerate solve_tril with TMA (#26746)
Jiangyun Zhu committed
A
[LoRA] LoRA cuda graph specialization (#25914)
Andy Lo committed
Y
[Model][VLM] Support Bee-8B Model (#27012)
Yi Zhang committed
October 19, 2025
Y
Fix typo in ValueError message: use `kv_role` instead of `kv_disagg_role` (#27166)
Yongtao Huang committed
S
[Bugfix] Fix error with penalties when speculative decoding and structural output are enabled (#26586)
Sergei Skvortsov committed
C
[Misc] Move utils to avoid conflicts with stdlib, and move tests (#27169)
Cyrus Leung committed
I
[Chore] Separate out `vllm.utils.network_utils` (#27164)
iAmir97 committed
J
output type conversion fix (#27159)
Jianyu Huang committed
C
[Benchmark] Convenience script for multiple parameter combinations (#27085)
Cyrus Leung committed
D
[Chore] Separate out hashing utilities from vllm.utils (#27151)
dongbo910220 committed
2
[BugFix] Fix lazy imports involving outlines_core (#27158)
22quinn committed