Junglewise Threat Intelligence

CVE-2026-88047: Tesseract stack buffer overflow in ReadNormProtos

CVE-2026-88047 · Severity: high · CVSS 7.8 · Published 2026-09-10

Technologies: Tesseract OCR Tesseract. Vendors: Tesseract OCR.

Executive brief

Tesseract is an open-source OCR (optical character recognition) engine used to extract text from images and documents. A vulnerability in its model-loading code allows attackers to crash the application or potentially execute arbitrary code by providing a malicious trained model file (.traineddata). This affects any application using Tesseract to process untrusted document models.

Technical details

The vulnerability is a stack buffer overflow (CWE-121) in Classify::ReadNormProtos (src/classify/normmatch.cpp:193) where the NORMPROTO component of a .traineddata file is parsed. The function uses std::istream::operator>>(char*) to extract a whitespace-delimited token into a fixed 61-byte stack buffer without setting a stream width limit. Since the line buffer allows up to 99 characters and the destination buffer is only 61 bytes, a token longer than 60 characters causes up to 39 bytes of attacker-controlled data to overflow the stack buffer during TessBaseAPI::Init. This corrupts adjacent stack memory (return addresses, frame pointers, adjacent variables), resulting in denial of service or potential control-flow hijacking. Builds using Apple's libc++ C++20 bounded array overload are incidentally protected; typical libstdc++ builds remain vulnerable. The fix (applied in commit 1bda507, patched in version 5.5.4) adds std::setw() to bound the extraction to the buffer size.

Affected products

  • Tesseract OCR Tesseract 5.5.3 and earlier

Timeline

  • 2026-08-31: disclosed
  • 2026-08-25: patched: Commit 1bda507 addresses the vulnerability
  • 2026-09-10: advisory: CVE-2026-88047 published

References

Related threats