[Test][Misc] Move 2 size cases to double_node (#15141)
### What this PR does / why we need it? This PR updates the test configuration for `DeepSeek-V4-Pro-w4a8-PD-no-prefix` by reducing the `--max-num-batched-tokens` from `8192` to `4096` to optimize memory usage or adjust test sizes. Additionally, it introduces a new multi-node end-to-end test configuration for `GLM-5.2-W8A8C8-A3-Memcache-PD` with disaggregated prefill routing and memcache KV pool. ### Does this PR introduce _any_ user-facing change? No ### How was this patch tested? By running weekly test. - vLLM version: v0.27.1 - vLLM main: https://github.com/vllm-project/vllm/commit/ba07e4a48fc951300d97eb506217dd530583dea3 --------- Signed-off-by: chen-commits <1636718796@qq.com> Signed-off-by: chen <1636718796@qq.com>
C
chen-commits committed
f8f130809358e2e4451a0b9ed6ec82902f8d50a0
Parent: 4b25eb3
Committed by GitHub <noreply@github.com>
on 8/29/2026, 6:28:43 AM