Cybersecurity experts are raising concerns that the stringent guardrails imposed by leading artificial intelligence developers, including OpenAI and Anthropic, are significantly impeding the crucial work of offensive cybersecurity researchers. These researchers, tasked with identifying unknown system vulnerabilities and developing methods to exploit them before malicious actors can, argue that current AI model restrictions are making their jobs considerably more difficult, potentially leaving businesses and governments more exposed to cyber threats.
The issue stems from AI companies' efforts to prevent their powerful models from being used for nefarious purposes, such as creating sophisticated cyberattacks. This has led to the implementation of strict usage policies and, in some cases, specialised 'vetted access' programmes like OpenAI's Trusted Access for Cyber and Anthropic's Cyber Verification Program. However, researchers contend that the very tools designed to prevent misuse are simultaneously blocking legitimate security testing.
Mark Dowd, a prominent security researcher known for finding and selling 'zero-day' vulnerabilities to Western governments, voiced his discomfort on a recent cybersecurity podcast, stating, “it’s not really comfortable to me that these random large companies are making arbitrary decisions about what is safe in security and what’s not.” Dowd's work involves proactively probing systems for weaknesses, a process that increasingly relies on AI tools.
Chris Anley, chief scientist at security consulting giant NCC Group, highlighted the dual nature of AI tools in cybersecurity. He explained that asking an AI model to attempt to exploit a bug is a key step in confirming a vulnerability's existence and severity. Yet, if AI guardrails prevent the model from answering such a prompt, it directly hinders defensive efforts. Anley likened it to a hammer: “You can’t build a house without a hammer. It’s definitely a tool but it’s also irreducibly a weapon as well.” When encountering these roadblocks, Anley and his colleagues sometimes resort to open-source AI models that lack such restrictions.
The regulatory landscape adds another layer of complexity. While the UK ICO focuses on data protection and responsible AI use, and the EU AI Act aims to classify and regulate AI systems based on risk, neither explicitly addresses this specific tension between guardrails and offensive cybersecurity research. The current situation suggests a need for a nuanced approach that allows legitimate security testing while mitigating the risks of misuse. For now, some researchers, like Paolo Stagno, CTO at Crowdfense, describe AI companies' vetted programmes as treating customers “like children who need babysitting,” leading them to avoid using cloud-based AI for vulnerability discovery due to the risk of sensitive data leakage.
The implications for UK businesses and the broader economy are significant. If offensive cybersecurity researchers are unable to effectively utilise advanced AI tools to uncover vulnerabilities, critical weaknesses could remain undiscovered for longer. This increases the window of opportunity for malicious hackers, potentially leading to more frequent and severe cyberattacks, data breaches, and financial losses for UK organisations. The balance between AI safety and cybersecurity innovation is a delicate one that requires ongoing dialogue and potentially new regulatory frameworks to ensure both security and progress.