Executive brief
Hugging Face Datasets is a library used by developers and researchers to download and manage large machine learning datasets. A security flaw allows a local attacker to trick the library into writing files to unauthorized locations on the computer's hard drive. In shared computing environments, such as multi-user servers or shared research workstations, this could allow an attacker to overwrite sensitive system files, potentially leading to a full system takeover or data loss.
Technical details
A symlink-following vulnerability exists in the `Extractor.extract()` function within `src/datasets/utils/extract.py`. The root cause is a reliance on `shutil.rmtree(output_path, ignore_errors=True)` to clear the extraction directory; in Python 3.8+, this function fails silently when the path is a symbolic link. Because the extraction output paths are predictable (based on a non-salted SHA256 hash of the source URL), a local attacker can pre-plant a symlink at the expected path. When a victim subsequently triggers an extraction, the library follows the symlink and writes the archive contents to the attacker's chosen target directory. This bypasses existing 'Zip-Slip' mitigations (safemembers) which only validate archive member names relative to the resolved output path. The issue is fixed in commit ad2d853 by explicitly unlinking the path if it is a symbolic link before extraction.
Affected products
- Hugging Face datasets <= 5.00
Timeline
- 2026-07-03: disclosed: Issue reported to maintainers via GitHub
- 2026-07-06: patched: Fix merged into main branch via commit ad2d853
- 2026-07-23: advisory: CVE-2026-65010 published