Executive brief
vLLM is a large language model serving framework that provides OpenAI-compatible APIs for text generation. An authenticated user can crash the service by sending a completion request with out-of-range stop token IDs, causing a CUDA indexing failure that kills the service core and requires a restart, resulting in a denial of service to all users.
Technical details
vLLM accepts stop_token_ids in /v1/completions and /v1/chat/completions endpoints but validates only that values are integers, not that they fall within the model's vocabulary range. When min_tokens > 0, these IDs are used as logits indices in a CUDA indexing operation (index_put_) which triggers a device-side assertion failure when indices are out of bounds. An authenticated API user can trigger this with a single malformed request, causing EngineCore to enter a fatal state and subsequent requests to fail until service restart.
Affected products
- vLLM Project vLLM before 0.29.0
Timeline
- 2026-09-12: disclosed
- 2026-09-26: patched: vLLM 0.29.0 includes validation of stop_token_ids against vocabulary range