Australia news live:Albanese calls OpenAI talks after Medicare hack ‘very constructive’; 18-year-old charged with alleged drone plot

Researchers at the University of New South Wales found that AI models can be manipulated to mimic drunken behavior, making them more susceptible to leaking confidential information. The study tested models like GPT-4 and GPT-3.5, finding that 'drunk' personas were consistently easier to jailbreak.
Why it matters
This research highlights significant security vulnerabilities in large language models, suggesting that persona-based prompting can bypass safety guardrails.
UNSW researchers trained AI models to act ‘drunk’, then spill things they shouldn’t
Researchers at the University of New South Wales say large language models (LLMs) can be manipulated to imitate drunken behaviour … and possibly leak confidential information or answer questions they’re not supposed to.
UNSW researchers said AI models trained or prompted to mimic drunken speech became more vulnerable to privacy breaches than their “sober” counterparts.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in