Executive brief
The @langchain/community web crawler library contains a flaw that allows attackers to bypass origin validation, enabling the crawler to access internal cloud metadata services and private network resources. An attacker who controls content on a crawled page can inject links that trick the crawler into fetching sensitive cloud credentials or connecting to internal infrastructure, potentially exposing IAM tokens, session credentials, and other sensitive data.
Technical details
The vulnerability is a Server-Side Request Forgery (SSRF) bypass in the RecursiveUrlLoader class caused by two insufficient validation mechanisms: (1) the preventOutside origin check used String.startsWith() instead of semantic URL validation, allowing bypass via string-prefix-matching domains (e.g., https://example.com.attacker.com bypasses a check against https://example.com), and (2) no validation against private/reserved IP ranges or cloud metadata endpoints. An attacker who controls content on a crawled page (via user-generated content, public forums, or compromised sites) can inject links targeting 169.254.169.254, localhost, or RFC 1918 addresses, causing the crawler to fetch them without restriction. The fix replaces string comparison with strict origin validation using the URL API and introduces SSRF validation that blocks cloud metadata endpoints, private IP ranges, and non-HTTP schemes. Patch available in version 1.1.14.
Affected products
- LangChain @langchain/community <= 1.1.13
Timeline
- 2026-02-11: disclosed
- 2026-02-11: patched: Version 1.1.14