Executive brief
NVIDIA TensorRT-LLM, a library used to optimize and run large language models on Linux, contains a security flaw in how it handles model files. An attacker with local access to the system could provide a malicious model file that, when loaded, allows them to take control of the system. This could result in the theft of sensitive data, unauthorized changes to software, or a complete system takeover.
Technical details
A deserialization vulnerability (CWE-502) exists in NVIDIA TensorRT-LLM for Linux within the restricted unpickler component used during model weight loading. The flaw allows a local, unauthenticated attacker to trigger the deserialization of untrusted data by providing a specially crafted model file. Because the unpickler does not sufficiently restrict the classes it can instantiate, an attacker can achieve arbitrary code execution, escalate privileges, or access sensitive information. The vulnerability affects versions up to and including v1.3.0 rc14.
Affected products
- NVIDIA TensorRT-LLM v1.3.0 rc14 and earlier
Timeline
- 2026-07-14: disclosed: Initial publication of CVE-2026-24233