Summary
Issue 1: EXIF orientation not normalized → The image orientation processed by the model differs from how humans view it, introducing interpretation bias.
Issue 2: PNG tRNS not explicitly flattened before converting to RGB → After conversion, transparent/semi-transparent pixels are rendered unexpectedly, making otherwise subtle overlay elements visible and distorting the input content. (This attack is similar to AlphaDog: RGBA handling is already correct in vLLM, but since tRNS permits RGB images, the correct processing path isn’t taken.)
Issue 3 : Pillow only loads the first frame when loading APNG or GIF files.
Root Cause
- Rotation: After opening an image,
ImageOps.exif_transpose is not called to normalize EXIF orientation.
- Transparency: Only RGBA→RGB is flattened with a background; PNGs carrying
tRNS in P/L/RGB + tRNS and other non-RGBA modes take the image.convert("RGB") path, which implicitly discards/remaps transparency semantics.
Affected Code
|
def _convert_image_mode(self, image: Image.Image) -> Image.Image: |
|
"""Convert image mode with custom background color.""" |
|
if image.mode == self.image_mode: |
|
return image |
|
elif image.mode == "RGBA" and self.image_mode == "RGB": |
|
return rgba_to_rgb(image, self.rgba_background_color) |
|
else: |
|
return convert_image_mode(image, self.image_mode) |
|
def convert_image_mode(image: Image.Image, to_mode: str): |
|
if image.mode == to_mode: |
|
return image |
|
elif image.mode == "RGBA" and to_mode == "RGB": |
|
return rgba_to_rgb(image) |
|
else: |
|
return image.convert(to_mode) |
|
def rgba_to_rgb( |
|
image: Image.Image, |
|
background_color: tuple[int, int, int] | list[int] = (255, 255, 255), |
|
) -> Image.Image: |
|
"""Convert an RGBA image to RGB with filled background color.""" |
|
assert image.mode == "RGBA" |
|
converted = Image.new("RGB", image.size, background_color) |
|
converted.paste(image, mask=image.split()[3]) # 3 is the alpha channel |
|
return converted |
Current state: ImageOps.exif_transpose is not used. (Although the rescale_image_size function (https://github.com/vllm-project/vllm/blob/main/vllm/multimodal/image.py#L14) exists and includes a transpose parameter, I’ve found that it doesn’t seem to be called anywhere outside the test directory.)
Call order: _convert_image_mode runs first; if the conditions are met, convert_image_mode is called.
Issue: Only the “RGBA → RGB” path is explicitly flattened. P, L, or RGB with tRNS all fall back to image.convert("RGB"). For PNGs that include tRNS, convert("RGB") directly produces 24-bit RGB, leading to:
P mode: The transparent index becomes an actual RGB color (often black, white, or an undefined background), so transparency is lost.
L/LA and RGB + tRNS: convert("RGB") doesn’t composite against a chosen background first, so elements that relied on transparency to be hidden or softened become solid.
Impact & Scope
- Impact: Pixels the model sees can diverge from operator expectations (due to orientation or transparency handling), potentially altering downstream reasoning.
- Scope: The image I/O and mode-conversion paths in
vllm/multimodal/image.py. The existing RGBA→RGB flattening is correct; the issues center on missing EXIF normalization and non-RGBA tRNS not being explicitly composited.
Case
EXIF: http://qiniu.funxingzuo.top/exif_orient_180.jpg
tRNS: http://qiniu.funxingzuo.top/hello.png
Fix
A fix for this vulnerability was merged here: #44974
Summary
Issue 1: EXIF orientation not normalized → The image orientation processed by the model differs from how humans view it, introducing interpretation bias.
Issue 2: PNG tRNS not explicitly flattened before converting to RGB → After conversion, transparent/semi-transparent pixels are rendered unexpectedly, making otherwise subtle overlay elements visible and distorting the input content. (This attack is similar to AlphaDog: RGBA handling is already correct in vLLM, but since tRNS permits RGB images, the correct processing path isn’t taken.)
Issue 3 : Pillow only loads the first frame when loading APNG or GIF files.
Root Cause
ImageOps.exif_transposeis not called to normalize EXIF orientation.tRNSinP/L/RGB + tRNSand other non-RGBA modes take theimage.convert("RGB")path, which implicitly discards/remaps transparency semantics.Affected Code
vllm/vllm/multimodal/image.py
Lines 77 to 84 in 16b37f3
vllm/vllm/multimodal/image.py
Lines 37 to 43 in 16b37f3
vllm/vllm/multimodal/image.py
Lines 26 to 34 in 16b37f3
Impact & Scope
vllm/multimodal/image.py. The existing RGBA→RGB flattening is correct; the issues center on missing EXIF normalization and non-RGBAtRNSnot being explicitly composited.Case
EXIF: http://qiniu.funxingzuo.top/exif_orient_180.jpg
tRNS: http://qiniu.funxingzuo.top/hello.png
Fix
A fix for this vulnerability was merged here: #44974