Executive brief
The lxml_html_clean library, used for sanitizing HTML to prevent security risks, fails to properly scrub malicious JavaScript from certain link types. If an application is configured to allow custom attributes, an attacker could provide a specially crafted link that executes code in a user's browser when clicked. This could lead to unauthorized actions or data theft from users interacting with the affected website.
Technical details
A Cross-Site Scripting (XSS) vulnerability exists in lxml_html_clean due to an incomplete list of disallowed inputs in the link attribute allow-list. The 'Cleaner' class uses 'iterlinks()' to identify and scrub URL schemes, but 'iterlinks()' only checks attributes defined in 'lxml.html.defs.link_attrs', which lacks namespaced attributes like 'xlink:href'. When 'safe_attrs_only' is set to False, malicious 'javascript:' payloads in 'xlink:href' attributes (common in SVG and MathML) are bypassed. This allows for stored XSS if a victim clicks the affected link. The issue is patched in lxml_html_clean version 0.4.5.
Affected products
- fedora-python lxml_html_clean < 0.4.5
- lxml lxml <= 6.1.0
Timeline
- 2026-05-10: disclosed: Reported by Guillem Lefait
- 2026-06-01: advisory: Initial GitHub Advisory published
- 2026-07-08: patched: Advisory updated with patch information