Executive brief
NVIDIA Triton Inference Server is a widely-used inference platform that runs machine learning models in production. A vulnerability in the Linux version allows a remote attacker to cause excessive iteration loops, leading to denial of service and unavailability of the inference service for legitimate users and applications.
Technical details
The vulnerability involves a logic flaw in Triton Inference Server for Linux that permits an attacker to trigger excessive iteration in a processing loop without proper bounds checking or rate limiting. The attack is network-accessible and does not require authentication or user interaction. A successful exploit causes the server to consume CPU resources in an infinite or near-infinite loop, rendering the service unavailable. No patch details are currently available in the advisory, but NVIDIA has assigned CVE-2026-16497 and published a security bulletin.
Affected products
- NVIDIA Triton Inference Server <UNKNOWN>
Timeline
- 2026-09-08: disclosed