Executive brief
Picklescan, a tool used to detect malicious code in AI models, can be bypassed by specially crafted ZIP archives. By intentionally corrupting the internal checksum (CRC) of a ZIP file, an attacker can cause the scanner to skip the file entirely while the model remains loadable by frameworks like PyTorch. This allows malicious code to enter a production environment undetected, potentially leading to full system compromise when the model is executed.
Technical details
Picklescan fails to scan ZIP archives if they contain files with a mismatched Cyclic Redundancy Check (CRC). The tool relies on Python's built-in zipfile module, which raises a BadZipFile exception when encountering CRC errors, causing Picklescan to abort the scan for that archive without returning results. Because machine learning frameworks like PyTorch often bypass CRC checks during model loading, an attacker can craft a ZIP-archived model with an intentional CRC mismatch to evade detection while ensuring the malicious payload still executes upon loading. This vulnerability is addressed in version 0.0.31 by implementing a relaxed ZIP parsing mechanism that ignores CRC mismatches.
Affected products
- mmaitre314 picklescan <= 0.0.30
Timeline
- 2025-09-08: patched: Fix committed in version 0.0.31
- 2025-09-10: disclosed: Advisory published on GitHub
References
- https://github.com/mmaitre314/picklescan/security/advisories/GHSA-mjqp-26hc-grxg
- https://github.com/mmaitre314/picklescan/commit/28a7b4ef753466572bda3313737116eeb9b4e5c5
- https://github.com/mmaitre314/picklescan/blob/v0.0.29/src/picklescan/relaxed_zipfile.py
- https://huggingface.co/jinaai/jina-embeddings-v2-base-en/resolve/main/pytorch_model.bin?download=true
- https://huggingface.co/jinaai/jina-embeddings-v2-base-en/tree/main