Disclosure summary
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, Harmony tool continuations submitted through "POST /v1/responses" requests rebuild the next-turn engine input without preserving the cache_salt value, placing the continuation prefix in the global unsalted cache namespace even when the caller enabled salting. On deployments with prefix caching enabled, which is the default, an authenticated tenant who can reconstruct a victim's low-entropy post-tool history can submit the same continuation and use the cached_tokens_per_turn count to determine whether the prefix was previously processed, defeating the intended tenant isolation of salted prefix caching. This issue is fixed in version 0.30.0.
Source-reported weakness categories
CWE-200, CWE-524
Source-specific records & product guidance
Sources retain their own attribution and scoring. Follow the original record to confirm affected versions, fixed releases, and configuration conditions.
NIST National Vulnerability Database · NVD-CVE-2026-105752
Open original source · Updated Oct 06, 2026
Only CPE matches marked vulnerable=true are indexed. AND/OR platform conditions must be checked in the original NVD record.
GitHub Reviewed Security Advisories · GHSA-935w-9g4m-p28p
Open original source · Updated Oct 05, 2026
vLLM: Harmony tool continuations drop `cache_salt` — restoring a cross-tenant prefix-cache membership oracle
Source severity: LOW / 0
| Ecosystem | Package | Affected range | First patched |
|---|---|---|---|
| pip | vllm | < 0.30.0 | 0.30.0 |
Original records & references
PUBLISHED 2026-10-05T19:17:01-04:00
MODIFIED 2026-10-06T11:17:15-04:00
INGESTED 2026-10-06T11:45:40-04:00