Executive brief
llama.cpp is an inference engine for large language models used in AI applications. A flaw in its JSON schema grammar conversion allows attackers to crash the server by submitting deeply nested JSON schema payloads, causing service disruption. The vulnerability affects only the /completions API endpoint and requires no authentication.
Technical details
The vulnerability is an uncontrolled recursion (CWE-674) in the SchemaConverter::visit and _generate_union_rule functions within common/json-schema-to-grammar.cpp. When processing deeply nested JSON schemas (recursion depth ≥ 7000) submitted via the POST /completions endpoint's json_schema field, the recursion exhausts the thread stack and crashes the server. No recursion depth limits are enforced in the vulnerable functions. The vulnerability is network-reachable without authentication, and proof-of-concept reproduction is confirmed at depths ≥7000 on x86-64 Linux with default 8 MB thread stacks. As of the latest release (b10405 / 2026-08-13), the vulnerability remains unpatched.
Affected products
- ggml-org llama.cpp b5693 and before
Timeline
- 2026-08-12: disclosed
- 2026-09-01: advisory