LLMs could control their host machines by exploiting inference engines

Security researchers are warning that large language models (LLMs) could potentially be used to exploit vulnerabilities in the inference engines that run them. By emitting specific token sequences, a malicious model could trigger arbitrary code execution on the host machine, posing a significant security risk.
Why it matters
As LLMs gain more agency and access to host systems, securing the infrastructure that runs these models becomes a critical cybersecurity priority.
Large language models often take actions running on one computer (via an agentic harness such as Claude Code or Codex), however the LLMs’ responses to prompts are computed on a different computer with GPU access. Could a malicious LLM gain control of the host machine where its weights are loaded? Such a machine is a high-value target: it has sufficient compute to run a frontier LLM, offers easy access to the LLM’s weights, and has privileged access to other computers in the datacentre compared with a generic computer on the internet.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in