Disclosure summary
vLLM is an inference and serving engine for large language models. Prior to 0.28.0, the default mirrored multimodal LRU cache can commit a media hash in the frontend sender cache during multimodal rendering and before engine admission, while the engine receiver cache never receives the payload if that request is rejected. A later request reusing the same media hash causes MultiModalProcessorSenderCache to send no payload and MultiModalReceiverCache to reach an assertion with the message "Expected a cached item," producing a shared-service availability failure. This issue is fixed in version 0.28.0.
Source-reported weakness categories
CWE-617
Source-specific records & product guidance
Sources retain their own attribution and scoring. Follow the original record to confirm affected versions, fixed releases, and configuration conditions.
NIST National Vulnerability Database · NVD-CVE-2026-105753
Open original source · Updated Oct 06, 2026
Only CPE matches marked vulnerable=true are indexed. AND/OR platform conditions must be checked in the original NVD record.
GitHub Reviewed Security Advisories · GHSA-ph3r-5jfg-f84f
Open original source · Updated Oct 05, 2026
vLLM: Mirrored multimodal IPC caches desync after a rejected request — a later request reusing the same media hash trips a receiver assertion in the engine core
Source severity: MEDIUM / 0
| Ecosystem | Package | Affected range | First patched |
|---|---|---|---|
| pip | vllm | < 0.28.0 | 0.28.0 |
Original records & references
PUBLISHED 2026-10-05T19:17:01-04:00
MODIFIED 2026-10-06T11:19:48-04:00
INGESTED 2026-10-06T11:45:40-04:00