Executive brief
Ollama, a platform for running large language models, is vulnerable to a remote attack that can crash the entire server. By uploading a specially crafted, tiny model file (less than 1KB), an attacker can force the server to attempt to allocate massive amounts of memory, leading to an immediate and unrecoverable system failure. This disrupts all AI services provided by the server and may impact other applications running on the same host due to memory exhaustion.
Technical details
An uncontrolled memory allocation vulnerability (CWE-789) exists in Ollama's GGUF metadata parser within `fs/ggml/gguf.go`. The parser reads length and count fields (string lengths, tensor dimension counts, and metadata array counts) directly from GGUF files and uses them as allocation sizes for `make()` calls or slice bounds without validating them against the actual file size. A remote, unauthenticated attacker can trigger this by uploading a sub-1KB crafted GGUF file via the `/api/blobs` and `/api/create` endpoints, or by inducing a model pull. This results in Go runtime fatal errors (out-of-memory) or `makeslice` panics that bypass standard recovery middleware, causing a complete process crash. The vulnerability affects versions up to commit f0078ae.
Affected products
- Ollama Ollama up to commit f0078ae
Timeline
- 2026-06-03: disclosed: Reported via Huntr
- 2026-07-05: other: Public GitHub issue opened
- 2026-07-21: advisory: NVD publication date