Disclosure summary
vLLM is an inference and serving engine for large language models. Prior to 0.30.0, structured-output request failures can escape request-scoped validation and reach the EngineCore fatal-error path. A per-request backend mismatch can re-raise a grammar compilation exception, padding produced by the ngram_gpu speculative-decoding mode can pass a negative token to guidance validation, and the Rust frontend can admit empty structured-output values that the Python frontend rejects, allowing ordinary constrained-generation requests to terminate the shared engine. This issue is fixed in version 0.30.0.
Source-reported weakness categories
CWE-20, CWE-248, CWE-755
Source-specific records & product guidance
Sources retain their own attribution and scoring. Follow the original record to confirm affected versions, fixed releases, and configuration conditions.
NIST National Vulnerability Database · NVD-CVE-2026-105757
Open original source · Updated Oct 06, 2026
Only CPE matches marked vulnerable=true are indexed. AND/OR platform conditions must be checked in the original NVD record.
GitHub Reviewed Security Advisories · GHSA-85xf-c7hm-whqw
Open original source · Updated Oct 05, 2026
vLLM: Structured-output request errors escape the request boundary and terminate the shared EngineCore — engine-fatal denial of service (3 sites)
Source severity: MEDIUM / 0
| Ecosystem | Package | Affected range | First patched |
|---|---|---|---|
| pip | vllm | < 0.30.0 | 0.30.0 |
Original records & references
PUBLISHED 2026-10-05T19:17:02-04:00
MODIFIED 2026-10-06T11:19:48-04:00
INGESTED 2026-10-06T11:45:40-04:00