Disclosure summary
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, the /inference/v1/generate endpoint in the disaggregated scale-out path accepts caller-supplied tensors in the features.kwargs_data field, cache identifiers in the features.mm_hashes field, ranges in the features.mm_placeholders field, and wire-selected multimodal field processors without rebinding them to the active model renderer contract. Forged grid geometry, field types, or non-positive placeholder lengths can terminate the shared EngineCore; when an attacker knows or can induce a victim's content hash, forged cache hashes can poison or retrieve cross-request encoder-cache state; and dropped sparse placeholder masks can alter replayed transport semantics. This issue is fixed in version 0.30.0.
Source-reported weakness categories
CWE-20, CWE-617, CWE-639, CWE-668, CWE-704, CWE-1284
Source-specific records & product guidance
Sources retain their own attribution and scoring. Follow the original record to confirm affected versions, fixed releases, and configuration conditions.
NIST National Vulnerability Database · NVD-CVE-2026-105754
Open original source · Updated Oct 06, 2026
Only CPE matches marked vulnerable=true are indexed. AND/OR platform conditions must be checked in the original NVD record.
GitHub Reviewed Security Advisories · GHSA-ph72-cqr5-qpp7
Open original source · Updated Oct 05, 2026
vLLM: Scale-out disaggregated multimodal transport trusts caller-supplied features
Source severity: MEDIUM / 0
| Ecosystem | Package | Affected range | First patched |
|---|---|---|---|
| pip | vllm | < 0.30.0 | 0.30.0 |
Original records & references
PUBLISHED 2026-10-05T19:17:02-04:00
MODIFIED 2026-10-06T10:59:48-04:00
INGESTED 2026-10-06T11:45:40-04:00