Junglewise Threat Intelligence

CVE-2026-31253: Dao-AILab flash-attention insecure deserialization in checkpoint loading

CVE-2026-31253 · Severity: high · CVSS 7.3 · Published 2026-05-11

Vendors: PyPI.

Executive brief

Flash-attention is a library used to speed up the training and operation of large artificial intelligence models. A security flaw in how the software loads "checkpoint" files (saved states of a model) allows an attacker to execute malicious code on a user's system. This occurs if a user is tricked into loading a specially crafted checkpoint file during model training or evaluation, potentially leading to a full system compromise.

Technical details

An insecure deserialization vulnerability (CWE-502) exists in the flash-attention framework through commit e724e25. The load_checkpoint() function in checkpoint.py and loading logic in eval.py utilize torch.load() without the weights_only=True restriction. Because torch.load() uses the Python pickle module by default, it can be coerced into instantiating arbitrary Python objects. An attacker can exploit this by providing a maliciously crafted checkpoint file; when a victim attempts to load this file for model warmstarting or evaluation, the system executes arbitrary code embedded in the pickle stream. As of the advisory date, no patched version is specified for the affected 2.8.3 range.

Affected products

  • Dao-AILab flash_attn <= 2.8.3

Timeline

  • 2025-13-04: other: Vulnerable commit e724e2588cbe754beb97cf7c011b5e7e34119e62 identified
  • 2026-05-11: advisory: GitHub Advisory published
  • 2026-05-11: disclosed: NVD publication date

References

Related threats