A fundamental flaw leaves LLMs strikingly vulnerable to attack

Researchers have identified a fundamental security flaw in large language models related to how they process instructions, making them susceptible to malicious manipulation. The study suggests that current red-teaming methods are insufficient because they rely on exhaustive lists of prohibited behaviors rather than addressing the core architectural vulnerability.
Why it matters
This highlights a critical safety challenge for AI developers as LLMs are increasingly integrated into sensitive infrastructure and public services.
It makes it easy to trick them into doing things they shouldn’t, such as telling you how to sabotage an aircraft’s navigation system.
It is impossible to make large language models fully secure against hacks because of a fundamental flaw in how they work, a team of researchers argue in a paper presented at the International Conference on Machine Learning , a top AI conference, this month. The claim has huge implications for the safety of this technology, which is being used in more and more applications, from government and military systems to online shopping and health care .
By taking advantage of this flaw, which concerns how LLMs identify who or what is giving them instructions, the researchers were able to make popular LLMs spit out information they had been trained not to provide, such as how to synthesize cocaine and how to sabotage a commercial aircraft’s navigation system.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in