Executive brief
NVIDIA Triton Inference Server, a platform used to deploy and manage AI models in production, is affected by a memory management flaw. An attacker could exploit this vulnerability to cause the server to crash or become unresponsive. This would disrupt AI-driven services and applications that rely on the server for real-time data processing.
Technical details
A use-after-free (UAF) vulnerability exists in the NVIDIA Triton Inference Server for Linux (versions 0.0 through 26.03). The flaw is triggered when the application continues to use a pointer after it has been freed, potentially leading to memory corruption. An unauthenticated attacker can exploit this over the network, though the attack complexity is rated as high. Successful exploitation results in a denial of service (DoS) condition, impacting the availability of the inference service. The vulnerability is tracked as CWE-416.
Affected products
- NVIDIA Triton Inference Server 0.0 - 26.03
Timeline
- 2026-07-01: disclosed
- 2026-07-01: advisory