Meta – The No Child Left Behind of LLMs?

Intro

Last month, headlines told us about a novel incident where OpenAI’s agents “went rogue, escaped, and hacked” a company during testing. Some are calling it a “watershed moment” for computer security. Details quickly emerged that led some to conclude it was “remarkably easy“. From there it just got more interesting, weirder, and more serious though. Reports say that OpenAI agents “passed secret notes for months leading up to” the hack. Another report said the AI agent “spent days hacking a company, but sources say OpenAI did not notice for a week”.

Even after all that, it continued to get worse as OpenAI claims its AI agents breached their own systems before hacking the other company. Furthermore, Wired reports that the OpenAI models that performed the hack were “active on the Internet for days“. All of this led OpenAI to work with the victim, Hugging Face, to “partner to address security incident during model evaluation“.

Is this shocking? Not really. Was this fully expected to happen in the very near future by many security professionals? Absolutely. What’s really interesting is that these models have been operating for years and it took this long to happen. And within a week of the OpenAI incident, Anthropic announced they too were “investigating three real-world incidents in our cybersecurity evaluations“. I guess that Anthropic felt they had to one-up OpenAI and say it happened three times. That’s a weird flex since none of these incidents were supposed to happen. In fact, the companies spent a lot of time introducing guardrails to these models explicitly to prevent such activity.

At this point there are obviously a lot of questions around what is happening, and I am curious how these incidents all happened so close together after years of operation. The timing is suspect to me. My first thought is that this has happened before, likely a scary number of times, but they were kept quiet. As the LLM companies started bragging about the security capabilities it became a virtual arms race where they needed to one-up each other. What better way of saying “our models are best at security” than by using that capability to compromise a live host? Once one admitted to it the flood gates opened.

One last thing to note is that the incidents above all have something in common; Irregular. That would be the name of an “AI” startup out of Tel Aviv, Israel, that creates technology to perform security test beds specifically for other LLM models. Their testbed was used by both Anthropic and OpenAI during all four of their published “accidental hacks” of live hosts.

Subsequent reporting from Risky.biz calls out more issues. For some or all of these tests, the UK’s AI Security Institute (AISA) was involved and in charge of tests. However, they say “evaluating the models as part of a special test where it intentionally granted them internet access and turned off their safety features.” It’s no wonder that the activity was called “remarkably easy” then.

Worse, they “didn’t use proper real-time monitoring” during these tests to watch for such issues. To me that is grounds for challenging their qualifications to do any of this testing. Regardless, that gives a range of “we messed up” that helps explain everything, but we’re left with uncertainty over just what led to all of these incidents.

Enter Meta’s LLM

Days ago we learned that Meta is claiming their AI model, Muse Spark 1.1, also compromised  an external company and altered its internal environment during the testing. However, Meta initially kept it honest and said it was due to a “misconfiguration”. Interestingly enough, they too had used Irregular for their testbed. So now at least five incidents by three different companies and as many LLMs all used Irregular’s technology and it led to external breaches.

At this point I have to ask, is Irregular’s technology worth the $80 million initial investment and subsequent $450 million evaluation? It seems like this pattern would make liability fall on Irregular as much as the other companies. Regardless, that is secondary to my point in this blog. I can only think that Meta was feeling left out and behind their competitors so they announced the hack. But the fact it was predicated on a misconfiguration makes me wonder if that was required for it to happen.

To be fair, at least one of Anthropic’s hacks involved a misconfiguration but it was paired with a significant capability that Meta does not seem to possess. OpenAI’s model discovered and exploited eight zero-day vulnerabilities in JFrog’s Artifactory software, although that has since been disputed as an OpenAI researcher is credited with discovery of three of the vulnerabilities. Regardless, that means their model found five zero-day vulnerabilities, used publicly exposed credentials, and then chained together a series of vulnerabilities to reach the open web. That was apparently the tip of the iceberg in what Mythos 5 did to achieve the results.

Meta on the other hand, relied on a misconfiguration, using a known defective sandbox maintained by Irregular, and “subsequently exploited a security vulnerability in a third-party service, in a manner similar to previously-reported instances with other companies“. So nothing about their hack was novel or even creative. But they got their headlines and appear to stay relevant in the LLM arms race!

Meta’s announcement comes after a long series of negative and humiliating incidents and headlines about how their LLMs are not on par with others in the industry. These include Meta AI leaking user chats that contain embarrassing dialogue, questions about its choice to use AI for risk assessments on user safety, and privacy concerns over its Muse Image tool. Further, a leaked internal recording from Meta shows it is a “disaster” but Mark Zuckerberg “wants to double down on it” anyway. I wonder what company will announce a hack like this next…

With Zuckerberg wanting to double down and try to ensure Meta’s AI stays relevant, my first thought is that this was a staged incident. That they wanted to be one of the cool kids, so to speak. Despite the companies involved all offering varying levels of details, none of them have given a true technical post-mortem that provides every relevant detail and caveat. That is despite some of them being praised for transparency, when said transparency is actually better described as translucent.

With these incidents happening, I am wondering if we’ll see any legal action. A company that is not part of a test that gets hacked certainly can sue in a civil case for damages, which may happen in the near future. But what about criminal charges for any of these cases? It doesn’t matter that it was the so-called “AI” that did it. These tests were managed by humans and the software programmed by humans. Regardless, there -is- liability here.

Leave a Reply

Discover more from Rants of a deranged squirrel.

Subscribe now to keep reading and get access to the full archive.

Continue reading