CVE-2026-93436
vLLM through 0.29.0 fails to properly clean up decode-side metadata for rejected inference requests in prefill/decode disaggregated deployments. Remote attackers can submit requests with max_tokens=0 to exhaust decode-worker memory without bound until the worker restarts.
- Published Sep 17, 2026
- CVSS 8.7 high
- 0.8% chance of exploitation in the next 30 days (EPSS)
- A fix is available