Executive brief
vLLM is an open-source library used to run large language models efficiently. A flaw in how it processes images can cause the AI to misinterpret visual data, such as seeing rotated images or distorted colors where transparency was intended. This could allow an attacker to manipulate the model's output or bypass safety filters by providing specially crafted image files that the AI "sees" differently than a human would.
Technical details
A vulnerability exists in vLLM's image processing pipeline due to improper handling of EXIF orientation and PNG transparency (tRNS) data. Specifically, the library fails to call ImageOps.exif_transpose after opening images, and it only flattens transparency for RGBA modes, ignoring tRNS chunks in other modes like P, L, or RGB. This results in 'Misinterpretation of Input' (CWE-115), where the model processes distorted or incorrectly oriented pixels. An attacker can exploit this via the network by submitting crafted images to an inference endpoint to cause integrity failures or bypass vision-language model constraints.
Affected products
- vLLM Project vLLM >= 0.11.0, <= 0.23.0
- Red Hat Red Hat AI Inference Server 3
Timeline
- 2026-06-17: disclosed: Initial disclosure and CVE assignment
- 2026-06-18: advisory: GitHub advisory published and later withdrawn as a duplicate