Executive brief
CosyVoice, a large-scale voice generation and speech synthesis model, is vulnerable to a security flaw that allows for the execution of unauthorized commands. This occurs when the system loads a specially crafted AI model file from a directory specified by the user. If an attacker provides a malicious model file and a user attempts to start the voice server using that directory, the attacker can take full control of the victim's computer.
Technical details
An insecure deserialization vulnerability (CWE-502) exists in the gRPC server component of CosyVoice. The root cause is the use of the `torch.load()` function without the `weights_only=True` parameter when loading speech synthesis models from a user-specified directory. Because `torch.load()` utilizes the Python `pickle` module by default, it can be coerced into deserializing arbitrary Python objects. An attacker can exploit this by placing a malicious model file in a directory that the victim then points the gRPC server to during initialization. Successful exploitation results in arbitrary code execution with the privileges of the user running the server.
Affected products
- FunAudioLLM CosyVoice thru commit 6e01309e01bc93bbeb83bdd996b1182a81aaf11e
Timeline
- 2026-05-11: advisory: NVD publication date