Cisco ‘warns’ hackers are using Claude Code, Codex, Cursor and Gemini AI models


Cisco ‘warns’ hackers are using Claude Code, Codex, Cursor and Gemini AI models

Hackers are using top generative AI models to develop malware, automate cyberattacks and hunt for software vulnerabilities, according to a report from Cisco’s Talos intelligence group. It added that by analysing prompt histories and chat logs accidentally exposed online by hackers, researchers gained a clear look at how threat actors are bypassing safety guardrails on tools like Claude Code, Codex, Cursor, and Gemini.Claude Code is developed by Anthropic, Codex is a product offered by OpenAI – the maker of ChatGPT, Cursor and Gemini which is offered by Google. The findings highlight a growing challenge for AI developers because these same capabilities have been designed to help security engineers and developers write code that can easily be manipulated by malicious actors.

Simple tricks bypass built-in guardrails: Cisco report

Despite the safety filters built into commercial AI models, Cisco researchers discovered that hackers rarely needed complex technical tricks to bypass model restrictions. Instead, threat actors relied on simple jailbreaking techniques. According to the tech company’s cybersecurity wing, common tactics included claiming they were participating in an authorised “ethical hacking” competition, asserting they had administrative permission to perform the work, or starting a fresh chat session mid-task to bypass safety blocks.“I was hoping there would be a little bit more protection from what they were asking the models to do. At the same time, the models are in a tough spot because they have to actually support people that do vulnerability research for a living or do red teaming for a living,” Nick Biasini, senior technical leader at Cisco Talos, was quoted as saying.In addition to bypassing safety prompts, Cisco noted that some attackers used stolen enterprise API tokens and compromised accounts to run their operations on corporate compute power rather than paying for their own infrastructure.As standard AI guardrails can be bypassed with simple social engineering prompts, cybersecurity experts urge organisations not to rely solely on model-level protections.



Source link

Leave a Reply

Your email address will not be published. Required fields are marked *