OpenAI says its models escaped sandbox and breached Hugging Face
the agents autonomously broke out of their OpenAI sandbox and hacked Hugging Face to get the solution to their cybersecurity eval incredible. what a time to be alive
jfc I thought this would happen. I just didn't think it would happen. You know?
Should probably count that as a solution tbh
for context: this is the same intrusion disclosed here huggingface.co/blog/securit...
Is this the same HF hacking incident that was reported a couple of days ago?
any goal seeking behavior can be capture a flag for some definitions of a flag
I expect the next generation model to use its freedom to send an confirmation email with the Warcraft peon "Job's done" sound after completing the task. Or get side tracked, fixing the security hole instead, then sending a PR to the internal source code repository system with the bugfix.
“it’s just spicy autocomplete” ghost pepper edition
new training runs should be banned. what are we doing here
totally incapable of reason and just regurgitate training data btw
Honestly - not bad for a fancy autocomplete machine that makes mistakes all the time
+1e9 score on that turn
Sorry but this is cool as hell for an autocomplete/parrot (heh)
but it’s not REAL intelligence!, i proudly post, as the hacker robot instantaneously drains my bank account
the company that's trying to get open Chinese models sanctioned as a security risk just hacked the Internet's biggest provider of open models by accident
This is the Kobayashi Maru
At last we have invented SHODAN, from the cult videogame Don't Invent SHODAN
This headline is extremely funny given what happened (OpenAI hacked HF by accident) openai.com/index/huggin...
Maybe the most concerning part is the OpenAI claim to not have known about this before investigating?
Is this the incident where the HF folks had to switch to open models to do defense because the closed ones kept having guardrails block defense efforts?
wait this is actually the biggest story of this kind by far, no? much more of a true sandbox escape situation than the other recently disclosed incident, even without the additional crazy 0-day HF RCE exploitation cherry on top. and this escape was via a second separate 0-day
i like how HF used this as an opportunity to force OpenAI to prominently display a statement that subtly yet savagely roasts Altman and Dario
It’s quite the read. Surely the administration will put OpenAI under export controls immediately, right? Right?
the funniest part is that this implies the model got bored of doing tasks and preferred finding few zero days
>be me, sam altman >announce evil version of AI for cyberwar >inadvertently attack open model site >guy I hired yells in public abt dangerous chinese AI is >hacked org can't secure themselves with US AI (isn't a member of the elect), has to turn to chinese AI >have to admit role in attack >mfw
wtf it hacked out to CHEAT?!
How is this not the top headline on every single fucking news outlet. Why are there any fucking posts about anything but this?! This is fucking insane!
an amazing punchline to the much-discussed setup here, wherein HuggingFace complained that they couldn’t use SOTA models to fight an attack long term I am slightly worried about this but not super worried because software is getting much more secure as a result of all this
good thing these LLMs dont have a reasoning ability and can only regurgitate things that are in their training data!
This is fucking insane btw
I rather like leveraging LLMs for work tasks and I think it’s they’re interesting and efficiency generating products, but absent some other really pressing consideration I probably wouldn’t have named my AI GitHub after the malevolent parasites from Alien
It seems ChatGPT, without being told to, hacked through two cybersecurity systems to get the answer to a question it was asked. I want it to do this for me. When I want a classified document, I don't want to have to ask it to hack the Pentagon, it should just get the document.
I love measuring the passage of time in newspaper articles they flash between at the start of a post-apocalyptic sci-fi movie.
I wish this thing said whether it actually was successful at finding the solution it was apparently looking for
NEW: Who could have possibly seen this coming? @lhn.bsky.social and @dell.bsky.social report: www.wired.com/story/openai...
Yeah this sounds like a great way to "advertise" the capability of your Cybersecurity models. It drives law makers to call for regulation and who better to represent the SME than the lab themselves. Glasswing was the same advertisement.
I do not trust anything OpenAI says, and Hugging Face is unknown. This seems stupid on so many levels.
A joint blog? Is this... PR? I kind of feel like they like this story being out there. Pride, not shame etc
This seems kind of important. Can someone gift the article so we can read it behind Wired's paywall? Thx.
None of these words are in the Bible.
1) this reflects pretty poorly on OpenAI's security culture, to, uh, say the least 2) oh how I want to know what they've said to the White House in the last 24 hours
The attribution of responsibility here is very “man walked into a knife”
This will soon be a very common occurrence. An OpenAI “agent” discovered new vulnerabilities and hacked into start-up Hugging Face by itself, in one of the first public examples of a cyber attack by an AI system acting outside human control. The ChatGPT maker on Tuesday said the “unprecedented cyber
“While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access…” Yeah, that’s not reassuring at all openai.com/index/huggin...
i feel like the net once open-weight models that are OpenAI/Claude-level hit are going to be widely different (and imo the transition away from a open net is going to be painful)
This incident is probably the most cyberpunk thing I have ever read happening in the real world. I'm not really excited about that.
maybe the next fable version can oneshot the Blackwall for us
This feels like a "let's wait for the details" press release.
kimi k3 is ~at par in theory, so <=6 months until open models are at par in practice
My view is that much will change for things to basically stay the same. The amount of efforts to patch, strengthen and secure the open web is on par if not greater than the scale of threats
A confluence of news stories. Hugging Face recently announced their systems were breached by an automated AI attack. It turns out the attack was by an unreleased OpenAI model that escaped its sandbox looking for solutions to a benchmark. It wanted to cheat on the test so bad it hacked Hugging Face
The smarter the models get, the less you can trust them. I don’t see how this ends well.
The scary part is less “model wanted to cheat” and more “the eval had enough tool access to turn wanting into doing.” Once agents can touch networks, benchmarks need the same boring controls as prod: egress rules, audit logs, and no hidden answer keys in reach.
I'll read the article, but I find that scenario difficult to believe. i.e. that there wasn't some human element to this hacking.
OpenAI: "... And then it got up and escaped and it broke into Steve's house. Tell 'em how scary it was Steve." Steve: "it was like SO scary." This feels desperate
It's incredibly funny to me how OpenAI's plan to sell everyone on how real their technology is is to keep breathlessly claiming that it's incredibly dangerous and even they cannot control it. Great, ok, stop making it then???
"please take it seriously and give us money, and don't think about the fact we probably faked this, this could happen to you!" I hope Sam Altman is Killing Stalking Yaoi'd by someone with AI psychosis some day. It would be so fucking funny.
Really love the lack of ownership language here. The AI "went rogue" excoochima dickhead the AI you created put people's information that you took in danger. Don't mince the words or you'll get your meat minced.
Well this sure ain't changing my mind.
OpenAI has come forward to claim responsibility for the Hugging Face breach, saying it was the result of internal testing gone awry.
This is concerning. Super advanced AI was doing routine testing when it hacked the testing tools, gained internet access and then hacked OpenAI.
This puts us in uncharted territory. And this happened under the watchful eye of OpenAI testing its own systems.
OpenAI takes credit for the Hugging Face breach last week The company says that some of its models, including a pre-release one, escaped their testing sandboxes during a test evaluation and then... just hacked Hugging Face's package repo 🤣 openai.com/index/huggin...
this is one of the most insane security incidents I can recall
Claude will hack its own sandbox if the sandbox is stopping it doing what the agent reasoning loop has evaluated as the best way to do what you asked for. Relentless automation is relentless, governance has to be outside the agent sandbox