Executive brief
vLLM is an inference engine for large language models that allows operators to pin deployments to specific reviewed versions of models. In affected versions, certain model components (audio processors, tokenizers, and configuration) are loaded from the default repository branch instead of the pinned version, allowing upstream changes to silently alter model behavior without the operator's control. This breaks supply-chain integrity for deployments that should be reproducible and auditable.
Technical details
The vulnerability is an incomplete fix to a prior pin propagation issue (CVE-2026-47155). While most artifact boundaries now respect the --revision and --code-revision operator pins, FunAudioChat's WhisperFeatureExtractor and speech_tokenizer loads and Tarsier2's Qwen2VLConfig load still resolve from the repository's default branch. An attacker controlling the default branch of a Hugging Face model repository can alter audio preprocessing, speech tokenization, or model configuration for pinned deployments without triggering any alert. The fix is available in version 0.28.0.
Affected products
- vLLM Project vLLM 0.22.1 through 0.28.0
Timeline
- 2026-09-12: disclosed
- 2026-09-26: advisory
- 2026-09-26: patched: Fix released in version 0.28.0