Cisco ‘warns’ hackers are using Claude Code, Codex, Cursor and Gemini AI models
Cisco’s Talos intelligence group reports that hackers are increasingly using generative AI models like Claude, Codex, and Gemini to automate cyberattacks and develop malware. The report notes that attackers easily bypass safety guardrails using simple social engineering and jailbreaking techniques.
Why it matters
This highlights a critical vulnerability in the current AI landscape, where tools designed for productivity are being weaponized by threat actors to lower the barrier for sophisticated cybercrime.
Hackers are using top generative AI models to develop malware, automate cyberattacks and hunt for software vulnerabilities, according to a report from Cisco’s Talos intelligence group. It added that by analysing prompt histories and chat logs accidentally exposed online by hackers, researchers gained a clear look at how threat actors are bypassing safety guardrails on tools like Claude Code, Codex, Cursor, and Gemini.Claude Code is developed by Anthropic, Codex is a product offered by OpenAI – the maker of ChatGPT, Cursor and Gemini which is offered by Google. The findings highlight a growing challenge for AI developers because these same capabilities have been designed to help security engineers and developers write code that can easily be manipulated by malicious actors.Simple tricks bypass built-in guardrails: Cisco reportDespite the safety filters built into commercial AI models, Cisco researchers discovered that hackers rarely needed complex technical tricks to bypass model restrictions.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in