Executive brief
Crawlee, a library used for web scraping and browser automation, contains a vulnerability that allows it to be tricked into making unauthorized requests to internal network services. By providing a malicious sitemap or robots.txt file, an attacker can force the crawler to probe internal admin panels, cloud metadata services, or other private resources not intended for public access. In certain configurations, this could also be used to read local files or interact with internal databases like Redis, potentially leading to data exposure or service disruption.
Technical details
Crawlee for Python (versions 1.0.0 to 1.6.3) is vulnerable to a two-layer Server-Side Request Forgery (SSRF). The first layer involves a lack of host validation for URLs found in sitemaps or robots.txt files, allowing the crawler to be coerced into sending GET requests to internal HTTP endpoints. The second layer occurs because nested sitemap fetching bypassed Pydantic URL validation, allowing non-HTTP schemes (like file://, gopher://, or dict://) to be passed directly to the underlying HTTP client. This is particularly severe when using the CurlImpersonateHttpClient, as it supports libcurl's full protocol suite, potentially enabling local file disclosure or RCE via protocol smuggling (e.g., gopher to Redis). The issue is patched in version 1.7.0 by enforcing host checks and validating URL schemes at the HTTP client level.
Affected products
- Apify Crawlee for Python >= 1.0.0, < 1.7.0
Timeline
- 2026-05-12: patched: Version 1.7.0 released
- 2026-05-15: advisory: GitHub Security Advisory published
- 2026-06-10: disclosed: CVE published to NVD