Executive brief
Apache OpenNLP, a machine learning library used for processing natural language text, contains a security flaw in how it handles dictionary files. An attacker can provide a specially crafted dictionary file that, when loaded by the software, allows them to read sensitive files from the server or perform unauthorized network requests. This could lead to the exposure of private data or internal system information.
Technical details
The DictionaryEntryPersistor class in Apache OpenNLP initializes a static SAXParserFactory without enabling FEATURE_SECURE_PROCESSING or disabling DTD processing. When the create(InputStream, EntryInserter) method or the public Dictionary(InputStream) constructor is invoked, the underlying XMLReader remains configured to resolve external entities and DOCTYPE declarations. A remote attacker can exploit this by providing a malicious XML-based dictionary file (such as a stop-word list) containing crafted DOCTYPE declarations. Successful exploitation allows for local file disclosure via file:// protocols or Server-Side Request Forgery (SSRF) via http:// protocols. The issue has been addressed in versions 2.5.9 and 3.0.0-M3 by aligning the parser configuration with the project's secure XmlUtil helper.
Affected products
- Apache OpenNLP < 2.5.9, 3.0.0-M1, 3.0.0-M2
Timeline
- 2026-05-01: disclosed: Initial disclosure on oss-security mailing list
- 2026-05-04: advisory: NVD publication date