Junglewise Threat Intelligence

CVE-2026-34760: vLLM-project vLLM improper input validation in audio downmixing

CVE-2026-34760 · Severity: medium · CVSS 5.9 · Published 2026-04-02

Technologies: vLLM Project vllm, vllm (PyPI), Vllm. Vendors: vLLM Project, PyPI, Vllm.

Executive brief

vLLM is a high-performance engine used to run and serve large language models (LLMs), including those that process audio. A vulnerability in its audio processing component causes a discrepancy between how humans hear audio and how the AI model interprets it due to non-standard audio mixing. This could lead to the AI model misinterpreting audio commands or data, potentially affecting the integrity of automated decisions or transcriptions.

Technical details

vLLM (versions 0.5.5 to 0.17.x) utilized the Librosa library for audio processing, which defaulted to using 'numpy.mean' for mono downmixing (to_mono). This deviates from the international standard ITU-R BS.775-4, which specifies a weighted downmixing algorithm. This discrepancy creates a 'semantic gap' where the audio processed by the AI model (via Librosa) does not match the audio heard by humans. An attacker could potentially exploit this inconsistency to provide audio that is interpreted differently by the model than by a human reviewer. The issue was addressed in version 0.18.0 by removing the Librosa dependency and replacing it with pyav for audio processing.

Affected products

  • vllm-project vllm >= 0.5.5, < 0.18.0

Timeline

  • 2026-03-21: patched: Fix merged into main branch via PR 37058
  • 2026-04-02: disclosed: CVE-2026-34760 published

References

Related threats