Executive brief
DLR-RM stable-baselines3 is a reinforcement learning library that loads trained models and replay buffers from files. The library deserializes these files using Python's pickle without validation, allowing attackers who control model or buffer files to execute arbitrary code. This is a critical risk when models are shared or downloaded from untrusted sources, potentially leading to complete system compromise.
Technical details
Three code paths in stable-baselines3 (PPO.load, load_replay_buffer, VecNormalize.load) use unsafe deserialization (cloudpickle.loads and pickle.load) on user-supplied files without gating or validation. An attacker can craft malicious model zips or replay buffer pickle files that execute arbitrary commands during deserialization via the pickle REDUCE opcode. While PyTorch tensor loading in the same codebase uses weights_only=True hardening, the pickle/cloudpickle paths remain unprotected.
Affected products
- DLR-RM stable-baselines3 up to 2.9.0
Timeline
- 2026-09-20: disclosed: CVE-2026-94093 published on NVD
- 2026-08-29: other: Vulnerability reported via GitHub issue #2281