Junglewise Threat Intelligence

OpenCC out-of-bounds read in UTF-8 processing

Severity: low · CVSS 3.1 · Published 2026-03-29

Technologies: BYVoid Opencc. Vendors: npm.

Executive brief

OpenCC is a library used to convert between different forms of Chinese characters. When processing malformed or truncated UTF-8 text, the library fails to validate input boundaries and reads past the end of the buffer, potentially exposing data in memory or causing the application to crash. This affects any application using OpenCC to handle user-supplied text.

Technical details

OpenCC versions before 1.2.0 contain out-of-bounds read vulnerabilities (CWE-125) in UTF-8 text processing. The root cause is a length validation failure: two code paths in MaxMatchSegmentation::Segment and Conversion::Convert fail to enforce the invariant that matchedLength ≤ remainingLength when handling malformed UTF-8 sequences. When processing truncated UTF-8 input (e.g., a 3-byte character missing its final byte), the library derives an incorrect length value, advances the input pointer beyond buffer bounds, and reads into adjacent memory—potentially beyond the null terminator. The attack vector is network-accessible if OpenCC processes untrusted UTF-8 input. No authentication or user interaction is required. Patch 1.2.0 fixes both issues by explicitly tracking input boundaries and clamping processed lengths to enforce the buffer invariant.

Affected products

  • BYVoid OpenCC < 1.2.0

Timeline

  • 2026-03-29: disclosed
  • 2026-03-27: patched: Version 1.2.0 released

References

Related threats