Executive brief
Apache OpenNLP is a Java library used for natural language processing tasks like tokenization, part-of-speech tagging, and model loading. A flaw in its model-loading mechanism allows an attacker who supplies a malicious model archive to trigger execution of arbitrary static code initializers or constructors from any class on the application classpath. If the classpath contains utility libraries with side effects during initialization (such as JNDI lookups or network calls), an attacker can exploit those to achieve data exfiltration, denial of service, or arbitrary command execution without requiring the victim to grant special permissions.
Technical details
The vulnerability exists in ExtensionLoader.instantiateExtension(Class, String), which loads a class by name via Class.forName() before checking whether it conforms to the expected extension interface (BaseToolFactory or ArtifactSerializer). Class.forName() with default semantics executes the target class's static initializer immediately, even if the type check that follows rejects the class as invalid. An attacker crafting a model archive can list any classpath class in the manifest.properties file, causing its static initializer to run during model load. While this is not direct remote code execution, exploitation is practical when the classpath includes libraries performing JNDI lookups, HTTP requests, file operations, or similar side effects during class initialization. A secondary narrower attack vector targets applications with BaseToolFactory or ArtifactSerializer subclasses that have side-effecting no-argument constructors. Patches (1.9.5, 2.5.9, 3.0.0-M3) introduce a package-prefix allowlist checked before Class.forName() is invoked.
Affected products
- Apache OpenNLP < 1.9.5, >= 2.0.0 and < 2.5.9, >= 3.0.0-M1 and < 3.0.0-M3
Timeline
- 2026-05-04: disclosed: Vulnerability published to GitHub Advisory Database and NVD
- 2026-05-04: patched: Patches released: 1.9.5, 2.5.9, 3.0.0-M3