Hi everyone!
Today in "sci-fi things that have happened": an #AI #agent that was being tested without guardrails to find #security vulnerabilities, found a vulnerability in the closed sandbox it was running in, escaped it's environs, then hacked into another website looking for the solution to the challenge it had been set...
Guardian article here, OpenAI explanation here, Simon Willison's write-up with much more technical detail here.
From a #cybersecurity point of view what's concerning about this is that OpenAI appears to have serious gaps in their network awareness and firewall configurations. If software on any of our servers started trying to hack another server anywhere else, alarm bells would go off all over the place.
Another item to note is that the victim of the attack, a site called Hugging Face, tried to use AI to analyse the attack and was prevented by guardrails. As Simon writes above this re-ignites the debate about guardrails: who gets to decide what they are, are they helping or hurting, etc. Clearly we need some, given the appalling use of poorly safeguarded models like Grok to generate deepfakes and worse, but it's clear that people who have access to AI without them, and who have tokens to burn, are in a position of considerable power.
At this point anyone who's read Neuromancer, arguably the first cyberpunk novel, will start worrying when Tessier-Ashpool SA will turn up, if they haven't already. The future's already here, etc. :-)
No comments:
Post a Comment