Junglewise Threat Intelligence

CVE-2026-93436: vLLM memory exhaustion in decode-side metadata cleanup

CVE-2026-93436 · Severity: high · CVSS 7.5 · Published 2026-09-17

Technologies: Vllm. Vendors: Vllm.

Executive brief

vLLM is a high-throughput inference engine for large language models (LLMs) commonly used to serve AI applications. In disaggregated deployments (where prefill and decode processing are split across separate workers), vLLM fails to clean up decode metadata when requests are rejected, allowing an attacker to exhaust decode-worker memory by submitting specially crafted requests, eventually forcing worker restarts and service disruption.

Technical details

This is a resource exhaustion vulnerability in vLLM's distributed inference system, specifically in the decode-side metadata cleanup logic for disaggregated prefill/decode deployments. The vulnerability occurs when inference requests with max_tokens=0 are submitted; these requests are rejected but the associated metadata is not properly cleaned up from the decode worker's memory. A remote attacker can repeatedly submit such requests over the network to unbounded exhaustion of heap memory, eventually triggering out-of-memory conditions and forcing decode worker restarts. The vulnerability affects vLLM through version 0.29.0, and patches should be available in later releases.

Affected products

  • vLLM vLLM through 0.29.0

Timeline

  • 2026-09-17: disclosed
  • other: CVE-2026-93436 assigned

References

Related threats